Skills

My technical background connects data analysis, machine learning, NLP research, rigorous evaluation, and research engineering for applied AI systems.


Technical Background

My technical background spans data analysis, machine learning, NLP, and research engineering. My previous experience as a data analyst provided a strong foundation in Python, SQL, R, statistical modelling, machine learning, and experimental design, together with practical experience working with real-world data and evaluating analytical results.

During my MSc in Language Technology, I extended this foundation into NLP research, working with PyTorch and Hugging Face on transformer-based models, machine translation, large language models, information retrieval, and model adaptation methods such as LoRA and Mixture-of-Experts.

My research places a strong emphasis on rigorous evaluation. I have experience with statistical testing, bootstrap confidence intervals, ablation studies, error analysis, and human evaluation, using these methods to understand model behaviour and assess the reliability of experimental findings.

More recently, my work has moved toward data-centric research, including corpus analysis, data filtering, and training-data attribution with Shapley values. From an engineering perspective, I use Git, Linux, Docker, and HPC/GPU environments to support reproducible research and large-scale experimentation.

Data Analysis & Visualization

Statistical analysis, exploratory data work, and visualization for real-world datasets, experimental results, and decision support.

PandasNumPyscikit-learnSQLRA/B TestingForecastingCustomer SegmentationPower BIStatistical AnalysisData VisualizationReporting

AI & Machine Learning

Deep learning, Transformer adaptation, fine-tuning workflows, and model evaluation for research and applied AI systems.

Machine LearningDeep LearningTransformersFine-tuningLoRA / PEFTModel Evaluation

NLP & Language Technology

Multilingual NLP, machine translation, retrieval, and language-technology systems for specialized domains.

Machine TranslationRAGInformation RetrievalLLM EvaluationMultilingual NLPTerminology Analysis

Responsible AI & Evaluation

Bias evaluation, human validation, statistical testing, and model behavior analysis for social and linguistic AI risks.

Bias EvaluationHuman AnnotationBootstrap TestingError AnalysisCohen's KappaVLM Framing

Data-Centric ML

Dataset diagnostics, filtering, attribution, and quality control for understanding how training data shapes model behavior.

Data FilteringData AttributionShapley AnalysisCorpus DiagnosticsQuality ControlData Auditing

ML Engineering

Research engineering skills for building reproducible model-training pipelines, experiments, demos, and deployments.

PyTorchHugging FacePythonDockerGitExperiment Tracking

Applied AI Systems

User-facing AI tools, dashboards, and interactive prototypes that turn model outputs into usable workflows.

GradioStreamlitDashboardsPlotlySQLiteUser-facing AI Tools