About Me
My name is Xiaojing Yang. I am a master's student in Language Technology at Uppsala University with a background in Business Analytics and four years of professional experience as a data analyst. My previous work gave me a strong foundation in data analysis, machine learning, statistical modelling, and experimental design.
My research focuses on multilingual NLP, model adaptation, and reliable language technology. It began with low-resource machine translation in the petroleum domain, where I explored parameter-efficient adaptation using LoRA. I later extended this work in my master's thesis to multilingual adaptation using Mixture-of-Experts and language-specific adapters. Through these projects, I became increasingly interested not only in how models can be adapted to improve their performance, but also in how training data shape model behaviour.
My recent work investigates this question through group-level training-data attribution using Shapley values. I also work on evidence-grounded retrieval, including GraphRAG for long documents, and bias evaluation in language and vision-language models. Across these areas, my broader goal is to develop NLP systems that are adaptable, interpretable, and reliable, together with methods for systematically auditing and understanding model behaviour.
Publications
I currently have one peer-reviewed publication, with several additional manuscripts under review and in preparation. My research outputs span multilingual NLP, machine translation, model adaptation, information retrieval, and data-centric evaluation, reflecting my broader interest in developing and understanding reliable language technologies.
LoRA Fine-Tuning of English-Norwegian NMT for the Oil & Gas Industry
Xiaojing Yang, Zhihan Li, Gege Sun, Mengyue Li, Meriem Beloucif
Published in the official EAMT 2026 proceedings. The ACL Anthology entry is forthcoming; to locate the paper in the proceedings PDF, search for the title "LoRA Fine-Tuning of English-Norwegian NMT for the Oil & Gas Industry".
EAMT 2026, Proceedings Volume 2
First author; lead experimental contributor
Modular Expert Architectures for Multilingual Domain Adaptation
Xiaojing Yang
Master's thesis on modular LoRA expert architectures for low-resource multilingual petroleum-domain NMT. Publicly summarized at a high level because the work forms the basis of an ongoing manuscript revision.
Beyond Routing: Diagnosing Modular LoRA Experts for Low-Resource Multilingual Petroleum-Domain Translation
Xiaojing Yang, Zhihan Li, Meriem Beloucif
Studies modular expert architectures for low-resource, domain-specific NMT, including language-specific LoRA adapters, learned routing, target-anchored synthetic data, terminology-aware evaluation, and routing analysis.
EMNLP 2026
First author; framework design and experimental lead
Projects
Selected research and technical projects in machine learning, NLP, retrieval, evaluation, and applied AI systems.

LoRA Fine-Tuning of English-Norwegian NMT for the Oil & Gas Industry
Published EAMT research on parameter-efficient English-Norwegian petroleum-domain NMT, showing how LoRA can adapt NLLB-200 with strong quality gains while updating less than 0.4% of model parameters.
Experience
Research, work, and teaching experience across AI, NLP, data analysis, and machine learning.
Research Projects in Language Technology
Designed and led research projects across multilingual NLP, machine translation, retrieval, data attribution, and bias evaluation.
Data Analyst
Completed various data analytics projects spanning customer intelligence, sales optimization, and predictive modeling across e-commerce, retail, and marketing domains.
Teaching Assistant, Machine Translation
Supported 25 students across assignment marking, lab sessions, and project supervision in a machine translation course.
Skills
My technical background connects data analysis, machine learning, NLP research, rigorous evaluation, and research engineering for applied AI systems.
Data Analysis & Visualization
Statistical analysis, exploratory data work, and visualization for real-world datasets, experimental results, and decision support.
AI & Machine Learning
Deep learning, Transformer adaptation, fine-tuning workflows, and model evaluation for research and applied AI systems.
NLP & Language Technology
Multilingual NLP, machine translation, retrieval, and language-technology systems for specialized domains.
Responsible AI & Evaluation
Bias evaluation, human validation, statistical testing, and model behavior analysis for social and linguistic AI risks.

