When Data Cleaning Becomes Bias: Target-Standard Specialization in Norwegian MT

Research
Responsible AI
Machine Translation
Evaluation
When Data Cleaning Becomes Bias: Target-Standard Specialization in Norwegian MT

Tech Stack

Python
PyTorch
Transformers
Hugging Face
NLLB-200
LoRA
SacreBLEU
chrF
Data Curation
Bootstrap
Statistics
SLIDE

Description

This project investigates whether data filtering changes the type of Norwegian produced by a machine translation model. The original petroleum corpus contains both Bokmål and Nynorsk. We examine what happens when the training data are filtered to contain mainly Bokmål.

As co-author and experimental analysis lead, I designed and conducted the controlled study of how Bokmål-oriented data filtering affects model behaviour and evaluation.

  • Designed size-controlled experiments and repeated them across three NLLB-200 model sizes.
  • Analyzed translation quality and written-standard output using statistical tests and human evaluation.

Project Highlights

Experimental Design

The study compares three training datasets:

Training datasetPurpose
Original mixed dataContains both Bokmål and Nynorsk
Bokmål-filtered dataContains mainly Bokmål
Same-size mixed subsetContains mixed data but has the same size as the filtered data

The filtered dataset and the same-size mixed dataset contain exactly the same number of sentence pairs. Comparing them allows us to separate the effect of data selection from the effect of data size.

Experimental Design figure

RQ1. Does filtering change the type of Norwegian produced by the model?

Answer: Yes. After training on Bokmål-filtered data, the model produces much more Bokmål and much less Nynorsk.

We compared two models on the same 1,313 test sentences: one trained on mixed Bokmål–Nynorsk data and one trained on Bokmål-filtered data. We then used SLIDE to identify whether each translation was written in Bokmål or Nynorsk. With mixed training data, 79.0% of the translations were Bokmål and 14.2% were Nynorsk. After Bokmål filtering, 93.4% were Bokmål and only 0.7% were Nynorsk. Human evaluation confirmed that the filtered model produced more Bokmål, but its translations were not clearly more accurate.

RQ1. Does filtering change the type of Norwegian produced by the model? figure

RQ2. Is the change caused by using less training data?

Answer: No. The change is caused by the type of examples selected.

The Bokmål-filtered model and the same-size mixed-data model were trained on exactly the same number of sentence pairs. However, the filtered model still produced much more Bokmål. This means that the model changed because Bokmål-oriented examples were selected—not because the amount of training data was reduced. The result was also consistent across the 600M, 1.3B, and 3.3B NLLB-200 models.

RQ2. Is the change caused by using less training data? figure

RQ3. Does the test set change which model looks better?

Answer: Yes. The Bokmål-filtered model performs better on the Bokmål test set, while the Original-subsampled model performs better on the mixed-standard test set. The test set can change which model looks better, so a higher BLEU score does not always mean a better translation.

RQ3. Does the test set change which model looks better? figure