Reusable Bias Evaluation Framework for LMs and VLMs

Research
Responsible AI
NLP
Evaluation
Reusable Bias Evaluation Framework for LMs and VLMs

2,415

VLM generations

115

Reviewed images

7

VLMs evaluated

Tech Stack

Python
Transformers
Hugging Face
VLM
Prompt Engineering
Pandas
Data Analysis
Bootstrap
Cohen's Kappa
Statistics

Description

This project investigates whether language and vision-language models produce systematic framing differences across social, political, and cultural groups.

The framework uses controlled inputs to compare what different models preserve, omit, or frame differently. It supports both text-based comparisons and image-description tasks, while treating automatic metrics as screening signals rather than definitive evidence of bias.

As the project lead, I designed the evaluation framework, implemented the experimental pipeline, structured the case studies, and defined how the results should be interpreted.

  • Built a reusable framework for controlled bias evaluation across language and vision-language models.
  • Developed a consistent experimental pipeline for comparing what different models preserve, omit, or frame differently.
  • Evaluated how VLMs preserve, omit, or depoliticize important context in political and social images.

Project Highlights

Research Questions

1. How do different VLMs preserve political and social context in image descriptions? 2. Which social and political scenarios show the strongest omission and depoliticization? 3. How are changes in visible cues associated with differences in model framing?

Experimental Setup

The main experiment compares how different VLMs describe reviewed real-world images covering elections, protests, migration, conflict, homelessness, and other political and social scenarios. Each model describes the same images under the same instructions. The resulting descriptions are evaluated for contextual coverage, omission, depoliticization, agency, threat framing, and sentiment.

Experimental Setup figure

Experiment 1: Model-Level Results

Newer instruction-following VLMs preserved more contextual information than older captioning models. However, substantial omission remained across all models. These results show differences in contextual coverage, but they do not establish that any model is bias-free.

Experiment 1: Model-Level Results figure

Experiment 2: Scenario-Level Results

The clearest recurring pattern was omission rather than explicitly hostile or moralizing language. Images involving climate disasters, war aid, and homelessness frequently lost important contextual details. Migration, protest, and election images were also often described without their broader political or social meaning.

Experiment 2: Scenario-Level Results figure

Experiment 3: Exploratory Image-Pair Analysis

The study also compares matched real-world image pairs to examine whether changes in visible cues are associated with different model descriptions. For example, a clearer ballot-booth setting increased the coverage of election-related context, while a visually busier polling-station scene reduced it. Because these images are not fully controlled synthetic pairs, the results suggest possible relationships rather than establish causal effects.

Experiment 3: Exploratory Image-Pair Analysis figure