Publications
Published, submitted, and planned research outputs in machine learning, NLP, retrieval, and model evaluation.
Parameter-Efficient Neural Machine Translation for Low-Resource Language Pairs: English-Norwegian in the Oil & Gas Domain
Xiaojing Yang, Zhihan Li, Gege Sun, Mengyue Li, Meriem Beloucif
Investigates parameter-efficient adaptation of neural machine translation for the Norwegian oil and gas domain, covering LoRA fine-tuning, hyperparameter optimisation, data-quality assessment, and model-scale evaluation.
EAMT
First author; lead experimental contributor
Beyond Routing: Diagnosing Modular LoRA Experts for Low-Resource Multilingual Petroleum-Domain Translation
Xiaojing Yang, Zhihan Li, Meriem Beloucif
Studies modular expert architectures for low-resource, domain-specific NMT, including language-specific LoRA adapters, learned routing, target-anchored synthetic data, terminology-aware evaluation, and routing analysis.
EMNLP 2026
First author; framework design and experimental lead
Structure-Aware Graph Retrieval for Evidence Grounding over Long Annual Reports
Xiaojing Yang, Zhihan Li, Meriem Beloucif
Evaluates how document structure, graph-based candidate expansion, and deterministic routing can improve evidence grounding over long financial reports, with robustness testing and held-out-year validation.
EMNLP 2026
First author; retrieval framework and evaluation lead
Group-level Training Data Attribution with Exact Shapley Analysis: A Controlled Machine Translation Case Study
Xiaojing Yang, Zhihan Li
Investigates group-level training-data attribution in neural machine translation using exact Shapley analysis, coalition training, size-matched baselines, bootstrap confidence intervals, and group-level statistical tests.
TBD
First author; attribution framework and experiment lead
Target-Standard Bias from Data Filtering in Norwegian Machine Translation
Zhihan Li, Xiaojing Yang
Studies how target-side data filtering affects written-standard preferences in English-Norwegian machine translation using controlled LoRA-based NLLB experiments and paired statistical validation.
WMT
Co-author; controlled experiment design and analysis lead
A Human-Validated Benchmark for Long-Form Financial Document Question Answering
Xiaojing Yang, Zhihan Li, Meriem Beloucif
Introduces a human-validated benchmark for long-form financial document question answering over annual reports, including evidence-object construction, retrieval evaluation, reranking, and statistical analysis.
TBD
First author; benchmark construction and retrieval evaluation lead