Target-Standard Bias from Data Filtering in Norwegian MT

Research
Responsible AI
Machine Translation
Evaluation
Target-Standard Bias from Data Filtering in Norwegian MT

Tech Stack

Python
PyTorch
Transformers
Hugging Face
NLLB-200
LoRA
SacreBLEU
chrF
Data Curation
Bootstrap
Statistics
SLIDE

Description

This project studies target-standard bias in Norwegian machine translation. The core question is whether target-side filtering toward Bokmal changes both model behavior and the way automatic metrics reward that behavior.

The method compares original, filtered, and size-controlled training conditions across NLLB-200 model scales. It combines automatic MT metrics, terminology evaluation, written-standard identification, out-of-domain FLORES checks, and diagnostic human assessment.

My contribution was to connect data filtering with responsible evaluation: I helped frame filtering as an auditable source of target-standard specialization, designed the size-controlled comparison, analyzed written-standard output shifts, and translated the result into practical safeguards for MT evaluation.

  • Designed a size-controlled comparison between original mixed-standard data and Bokmal-filtered data.
  • Evaluated translation quality, terminology behavior, written-standard output rates, and robustness across model scales.
  • Showed why reference-based metrics can encode target-standard preferences in multi-standard languages.
  • Added human-evaluation and deployment-interpretation framing to avoid treating all specialization as either good or bad.
  • Kept the project summary high-level until the manuscript path is settled.

Project Highlights

Written-Standard Data Diagnostics

The project treats filtering as a modeling decision, not a neutral preprocessing step, and evaluates how written-standard distributions affect MT scores and output behavior.

/projects/target-standard-bias/random-baseline-barchart.png