Projects
I have hands-on experience taking projects from idea and experimentation to implementation and evaluation, covering both research projects and applied AI systems.


Beyond Routing: Diagnosing Modular LoRA Experts for Low-Resource Multilingual Petroleum-Domain Translation
A diagnostic study of modular LoRA experts for low-resource petroleum translation.

FinRAG-Equinor: A Reliability-Audited Benchmark for Evidence Localization in Longitudinal Annual Reports
A reliability-audited benchmark for evidence-grounded RAG on 15 long annual reports, with 720 QA items and traceable report-, page-, and object-level evidence.

When Does GraphRAG Help? A Controlled Study of Evidence Localization in Long Annual Reports
A controlled study of when graph structure improves evidence localization in long annual reports, separating object recovery, page coverage, and candidate competition.
85.9%
Object Recall@10
91.8%
Page Recall@10
48,292
Graph nodes
When Data Cleaning Becomes Bias: Target-Standard Specialization in Norwegian MT
A controlled English-Norwegian MT study showing how Bokmål filtering creates target-standard specialization—and how reference choice can turn that specialization into an evaluation-bias problem.
Reusable Bias Evaluation Framework for LMs and VLMs
A reusable evaluation framework for matched-prompt and image-instruction bias studies, covering geographic, gender-occupation, and political/moral VLM framing cases.
2,415
VLM generations
115
Reviewed images
7
VLMs evaluated
E-commerce Sales Analytics: From Raw Transactions to Reliable Business Metrics
An analysis of 99K+ orders answering how sales changed, which categories and sellers drive value, and whether recorded payments can be trusted.
99,441
Orders
6
Source tables
99.58%
Payment match
Marketing Conversion A/B Testing: Does Advertising Increase Conversion?
A 588K-user experiment analysis asking whether ads improve conversion, how large the lift is, and whether the evidence is decision-ready.
588,101
Users
+0.769 pp
Absolute lift
1.71e-13
p-value


Beyond Routing: Diagnosing Modular LoRA Experts for Low-Resource Multilingual Petroleum-Domain Translation
A diagnostic study of modular LoRA experts for low-resource petroleum translation.

FinRAG-Equinor: A Reliability-Audited Benchmark for Evidence Localization in Longitudinal Annual Reports
A reliability-audited benchmark for evidence-grounded RAG on 15 long annual reports, with 720 QA items and traceable report-, page-, and object-level evidence.

When Does GraphRAG Help? A Controlled Study of Evidence Localization in Long Annual Reports
A controlled study of when graph structure improves evidence localization in long annual reports, separating object recovery, page coverage, and candidate competition.
85.9%
Object Recall@10
91.8%
Page Recall@10
48,292
Graph nodes
When Data Cleaning Becomes Bias: Target-Standard Specialization in Norwegian MT
A controlled English-Norwegian MT study showing how Bokmål filtering creates target-standard specialization—and how reference choice can turn that specialization into an evaluation-bias problem.
Reusable Bias Evaluation Framework for LMs and VLMs
A reusable evaluation framework for matched-prompt and image-instruction bias studies, covering geographic, gender-occupation, and political/moral VLM framing cases.
2,415
VLM generations
115
Reviewed images
7
VLMs evaluated
E-commerce Sales Analytics: From Raw Transactions to Reliable Business Metrics
An analysis of 99K+ orders answering how sales changed, which categories and sellers drive value, and whether recorded payments can be trusted.
99,441
Orders
6
Source tables
99.58%
Payment match
Marketing Conversion A/B Testing: Does Advertising Increase Conversion?
A 588K-user experiment analysis asking whether ads improve conversion, how large the lift is, and whether the evidence is decision-ready.
588,101
Users
+0.769 pp
Absolute lift
1.71e-13
p-value