When Data Cleaning Becomes Bias: Target-Standard Specialization in Norwegian MT
Tech Stack
Description
This project investigates whether data filtering changes the type of Norwegian produced by a machine translation model. The original petroleum corpus contains both Bokmål and Nynorsk. We examine what happens when the training data are filtered to contain mainly Bokmål.
As co-author and experimental analysis lead, I designed and conducted the controlled study of how Bokmål-oriented data filtering affects model behaviour and evaluation.
- Designed size-controlled experiments and repeated them across three NLLB-200 model sizes.
- Analyzed translation quality and written-standard output using statistical tests and human evaluation.
Project Highlights
Experimental Design
The study compares three training datasets:
| Training dataset | Purpose |
|---|---|
| Original mixed data | Contains both Bokmål and Nynorsk |
| Bokmål-filtered data | Contains mainly Bokmål |
| Same-size mixed subset | Contains mixed data but has the same size as the filtered data |
The filtered dataset and the same-size mixed dataset contain exactly the same number of sentence pairs. Comparing them allows us to separate the effect of data selection from the effect of data size.
RQ1. Does filtering change the type of Norwegian produced by the model?
Answer: Yes. After training on Bokmål-filtered data, the model produces much more Bokmål and much less Nynorsk.
We compared two models on the same 1,313 test sentences: one trained on mixed Bokmål–Nynorsk data and one trained on Bokmål-filtered data. We then used SLIDE to identify whether each translation was written in Bokmål or Nynorsk. With mixed training data, 79.0% of the translations were Bokmål and 14.2% were Nynorsk. After Bokmål filtering, 93.4% were Bokmål and only 0.7% were Nynorsk. Human evaluation confirmed that the filtered model produced more Bokmål, but its translations were not clearly more accurate.
RQ2. Is the change caused by using less training data?
Answer: No. The change is caused by the type of examples selected.
The Bokmål-filtered model and the same-size mixed-data model were trained on exactly the same number of sentence pairs. However, the filtered model still produced much more Bokmål. This means that the model changed because Bokmål-oriented examples were selected—not because the amount of training data was reduced. The result was also consistent across the 600M, 1.3B, and 3.3B NLLB-200 models.