Xiaojing Yang - AI Researcher and Machine Learning Engineer

Xiaojing Yang

AI Researcher & Machine Learning Engineer

NLP | Responsible AI | Data-Centric ML | Intelligent Systems

About Me

My name is Xiaojing Yang. I am a master's student in Language Technology at Uppsala University with a background in Business Analytics and four years of professional experience as a data analyst. My previous work gave me a strong foundation in data analysis, machine learning, statistical modelling, and experimental design.

My research focuses on multilingual NLP, model adaptation, and reliable language technology. It began with low-resource machine translation in the petroleum domain, where I explored parameter-efficient adaptation using LoRA. I later extended this work in my master's thesis to multilingual adaptation using Mixture-of-Experts and language-specific adapters. Through these projects, I became increasingly interested not only in how models can be adapted to improve their performance, but also in how training data shape model behaviour.

My recent work investigates this question through group-level training-data attribution using Shapley values. I also work on evidence-grounded retrieval, including GraphRAG for long documents, and bias evaluation in language and vision-language models. Across these areas, my broader goal is to develop NLP systems that are adaptable, interpretable, and reliable, together with methods for systematically auditing and understanding model behaviour.

Publications

I currently have one peer-reviewed publication, with several additional manuscripts under review and in preparation. My research outputs span multilingual NLP, machine translation, model adaptation, information retrieval, and data-centric evaluation, reflecting my broader interest in developing and understanding reliable language technologies.

PublishedJune 2026

LoRA Fine-Tuning of English-Norwegian NMT for the Oil & Gas Industry

Xiaojing Yang, Zhihan Li, Gege Sun, Mengyue Li, Meriem Beloucif

Pages 385-398. Tilburg, The Netherlands. European Association for Machine Translation. ISBN 9789403901411.

Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 1)

First author; lead experimental contributor

ThesisJune 2026

Modular Expert Architectures for Multilingual Domain Adaptation: Parameter-Efficient Norwegian Petroleum Translation with LoRA and Gated Routing

Xiaojing Yang

Master's thesis presenting a parameter-efficient multilingual adaptation framework that combines language-specific LoRA experts, gated routing, and target-anchored synthetic data for Norwegian petroleum-domain translation.

MSc Thesis, Uppsala University, 46 pages

Thesis author; framework design and experimental lead

Under ReviewOctober 2026

Beyond Routing: Diagnosing Modular LoRA Experts for Low-Resource Multilingual Petroleum-Domain Translation

Xiaojing Yang, Zhihan Li, Meriem Beloucif

Studies modular LoRA experts for low-resource petroleum NMT, showing that shared adaptation leads on controlled synthetic sources while learned-routing MoE generalises better to authentic source text.

EMNLP 2026

First author; framework design and experimental lead

Projects

I have hands-on experience taking projects from idea and experimentation to implementation and evaluation, covering both research projects and applied AI systems.

img
LoRA Fine-Tuning of English-Norwegian NMT for the Oil & Gas Industry

EAMT 2026 research on LoRA adaptation for low-resource petroleum-domain machine translation.

Research
Machine Translation
LoRA / PEFT
Low-Resource NLP
Domain Adaptation
Data Quality
Hyperparameter Optimization
Human Evaluation
img
Beyond Routing: Diagnosing Modular LoRA Experts for Low-Resource Multilingual Petroleum-Domain Translation

A diagnostic study of modular LoRA experts for low-resource petroleum translation.

Research
Machine Translation
Petroleum-Domain MT
Mixture-of-Experts
LoRA / PEFT
Routing Diagnostics
Low-Resource NLP
Evaluation
img
FinRAG-Equinor: Evidence-Grounded RAG Benchmark

A reliability-audited benchmark candidate for evidence-grounded RAG over 15 Equinor/Statoil annual reports, focusing on traceable report, page, and object-level evidence.

Research
RAG
Information Retrieval
Evaluation

Experience

Research, work, and teaching experience across AI, NLP, data analysis, and machine learning.

Research Projects in Language Technology

Uppsala University / Independent ResearchUppsala, Sweden
2024 - Present

Designed and led research projects across multilingual NLP, machine translation, retrieval, data attribution, and bias evaluation.

PythonPyTorch+9 more

Data Analyst

SAITISITE CHENGDU E-COMMERCE CO., LTDChengdu, China
2020 - 2024

Completed various data analytics projects spanning customer intelligence, sales optimization, and predictive modeling across e-commerce, retail, and marketing domains.

PythonPandas+10 more

Teaching Assistant, Machine Translation

Uppsala UniversityUppsala, Sweden
2026 - 2026

Supported 25 students across assignment marking, lab sessions, and project supervision in a machine translation course.

PythonPyTorch+5 more

Skills

My technical background connects data analysis, machine learning, NLP research, rigorous evaluation, and research engineering for applied AI systems.

Data Analysis & Visualization

Statistical analysis, exploratory data work, and visualization for real-world datasets, experimental results, and decision support.

PandasNumPyscikit-learnSQLRStatistical AnalysisData VisualizationReporting

AI & Machine Learning

Deep learning, Transformer adaptation, fine-tuning workflows, and model evaluation for research and applied AI systems.

Machine LearningDeep LearningTransformersFine-tuningLoRA / PEFTModel Evaluation

NLP & Language Technology

Multilingual NLP, machine translation, retrieval, and language-technology systems for specialized domains.

Machine TranslationRAGInformation RetrievalLLM EvaluationMultilingual NLPTerminology Analysis

Responsible AI & Evaluation

Bias evaluation, human validation, statistical testing, and model behavior analysis for social and linguistic AI risks.

Bias EvaluationHuman AnnotationBootstrap TestingError AnalysisCohen's KappaVLM Framing