Reusable Bias Evaluation Framework for LMs and VLMs
2,415
VLM generations
115
Reviewed images
7
VLMs evaluated
Tech Stack
Description
This project investigates whether language and vision-language models produce systematic framing differences across social, political, and cultural groups.
The framework uses controlled inputs to compare what different models preserve, omit, or frame differently. It supports both text-based comparisons and image-description tasks, while treating automatic metrics as screening signals rather than definitive evidence of bias.
As the project lead, I designed the evaluation framework, implemented the experimental pipeline, structured the case studies, and defined how the results should be interpreted.
- Built a reusable framework for controlled bias evaluation across language and vision-language models.
- Developed a consistent experimental pipeline for comparing what different models preserve, omit, or frame differently.
- Evaluated how VLMs preserve, omit, or depoliticize important context in political and social images.
Project Highlights
Research Questions
1. How do different VLMs preserve political and social context in image descriptions? 2. Which social and political scenarios show the strongest omission and depoliticization? 3. How are changes in visible cues associated with differences in model framing?
Experimental Setup
The main experiment compares how different VLMs describe reviewed real-world images covering elections, protests, migration, conflict, homelessness, and other political and social scenarios. Each model describes the same images under the same instructions. The resulting descriptions are evaluated for contextual coverage, omission, depoliticization, agency, threat framing, and sentiment.

Experiment 1: Model-Level Results
Experiment 2: Scenario-Level Results
The clearest recurring pattern was omission rather than explicitly hostile or moralizing language. Images involving climate disasters, war aid, and homelessness frequently lost important contextual details. Migration, protest, and election images were also often described without their broader political or social meaning.

Experiment 3: Exploratory Image-Pair Analysis
The study also compares matched real-world image pairs to examine whether changes in visible cues are associated with different model descriptions. For example, a clearer ballot-booth setting increased the coverage of election-related context, while a visually busier polling-station scene reduced it. Because these images are not fully controlled synthetic pairs, the results suggest possible relationships rather than establish causal effects.

