Projects

I have hands-on experience taking projects from idea and experimentation to implementation and evaluation, covering both research projects and applied AI systems.


img
LoRA Fine-Tuning of English-Norwegian NMT for the Oil & Gas Industry

EAMT 2026 research on LoRA adaptation for low-resource petroleum-domain machine translation.

Research
Machine Translation
LoRA / PEFT
Low-Resource NLP
Domain Adaptation
Data Quality
Hyperparameter Optimization
Human Evaluation
img
Beyond Routing: Diagnosing Modular LoRA Experts for Low-Resource Multilingual Petroleum-Domain Translation

A diagnostic study of modular LoRA experts for low-resource petroleum translation.

Research
Machine Translation
Petroleum-Domain MT
Mixture-of-Experts
LoRA / PEFT
Routing Diagnostics
Low-Resource NLP
Evaluation
img
Group-Level Training Data Attribution with Exact Shapley Analysis

Exact Shapley analysis of how training groups shape machine-translation behavior.

Research
Shapley Attribution
Training Data
Machine Translation
Data Auditing
Robustness Analysis
LoRA / PEFT
img
FinRAG-Equinor: A Reliability-Audited Benchmark for Evidence Localization in Longitudinal Annual Reports

A reliability-audited benchmark for evidence-grounded RAG on 15 long annual reports, with 720 QA items and traceable report-, page-, and object-level evidence.

Research
RAG
Information Retrieval
Evaluation
Data Quality
Human Evaluation
img
When Does GraphRAG Help? A Controlled Study of Evidence Localization in Long Annual Reports

A controlled study of when graph structure improves evidence localization in long annual reports, separating object recovery, page coverage, and candidate competition.

85.9%

Object Recall@10

91.8%

Page Recall@10

48,292

Graph nodes

Research
RAG
Information Retrieval
Evaluation
img
When Data Cleaning Becomes Bias: Target-Standard Specialization in Norwegian MT

A controlled English-Norwegian MT study showing how Bokmål filtering creates target-standard specialization—and how reference choice can turn that specialization into an evaluation-bias problem.

Research
Responsible AI
Machine Translation
Evaluation
img
Norwegian Petroleum Corpus

A reproducible pipeline for building English-Norwegian MT data and retrieval-ready text from public Equinor web pages and PDFs.

52,865

MT pairs

73,866

RAG chunks

0

Split leakage

Research
Data Quality
Data-Centric ML
Machine Translation
RAG
img
Replay-Based Continual Adaptation for English-Norwegian Petroleum Machine Translation

A continual LoRA adaptation study examining new-source learning, old-source forgetting, and replay-based retention in petroleum-domain MT.

Research
Machine Translation
LoRA / PEFT
Domain Adaptation
Evaluation
img
Reusable Bias Evaluation Framework for LMs and VLMs

A reusable evaluation framework for matched-prompt and image-instruction bias studies, covering geographic, gender-occupation, and political/moral VLM framing cases.

2,415

VLM generations

115

Reviewed images

7

VLMs evaluated

Research
Responsible AI
NLP
Evaluation
img
E-commerce Sales Analytics: From Raw Transactions to Reliable Business Metrics

An analysis of 99K+ orders answering how sales changed, which categories and sellers drive value, and whether recorded payments can be trusted.

99,441

Orders

6

Source tables

99.58%

Payment match

Data Science
Business Analytics
Data Quality
img
Customer Segmentation: Which Groups Behave Differently?

A five-segment customer analysis comparing K-Means, hierarchical clustering, Gaussian Mixture, and HDBSCAN with future-behavior validation.

4,335

Customers

0.974

Stability ARI

87.1%

Top repurchase

Data Science
Customer Analytics
AI/ML
img
Marketing Conversion A/B Testing: Does Advertising Increase Conversion?

A 588K-user experiment analysis asking whether ads improve conversion, how large the lift is, and whether the evidence is decision-ready.

588,101

Users

+0.769 pp

Absolute lift

1.71e-13

p-value

Data Science
Experimentation
Business Analytics
img
Retail Demand Forecasting: Which Model Generalizes Best?

A seven-model retail forecast comparing classical statistics, Gradient Boosting, Meta Prophet, and Google TimesFM through rolling validation.

3,049

Items in scope

5.89%

Rolling WAPE

5.75%

Final-test WAPE

Data Science
Forecasting
AI/ML
Business Analytics
img
90-Day Purchase Inactivity: Who Should Retention Prioritize?

A temporal customer-risk study using 48K monthly snapshots, purge-gap validation, four classifiers, and capacity-based targeting metrics.

48,079

Customer-months

0.628

Final PR-AUC

1.90×

Top-10% lift

Data Science
Customer Analytics
AI/ML