BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models

* Equally contributing first authors, ♠ Equally contributing second authors
1Institute of Artificial Intelligence, University of Central Florida, USA

Motivated by advancements in vision-language tasks by Large Multimodal Models (LMMs) and their limitations in producing unbiased responses, we present BBQ-V (BBQ-Vision), a multimodal extension of the BBQ dataset for assessing visual stereotype bias. Unlike prior benchmarks that lack diversity, rely on synthetic images, or use single-actor images, BBQ-V uses real-world, multi-actor images paired with deliberately ambiguous contexts and bias-probing questions. This enables more effective bias assessment, ultimately contributing to fairer and more inclusive LMMs that serve diverse communities equitably.

BBQ-V nine domains and 50 sub-domains

Figure: The BBQ-V benchmark spans nine diverse domains and 50 sub-domains to rigorously assess LMMs in visually grounded stereotypical scenarios. BBQ-V comprises over 14.1k carefully curated real-world, multi-actor image-question pairs built on 4,497 non-synthetic images.

Abstract

Stereotype biases in Large Multimodal Models (LMMs) perpetuate harmful societal prejudices, undermining the fairness and equity of AI applications. As LMMs grow increasingly influential, addressing and mitigating inherent biases related to stereotypes, harmful generations, and ambiguous assumptions in real-world scenarios has become essential. However, existing datasets evaluating stereotype biases in LMMs often lack diversity, rely on synthetic images, and frequently use single-actor images, leaving a gap in bias evaluation for real-world visual contexts. To address this gap, we introduce BBQ-Vision (BBQ-V), the most comprehensive framework for assessing stereotype biases across nine diverse categories and 50 sub-categories with real and multi-actor images. BBQ-V contains 14,144 image-question pairs and rigorously evaluates LMMs through carefully curated, visually grounded scenarios, challenging them to reason accurately about visual stereotypes. It offers a robust evaluation framework featuring real-world visual samples, image variations, and open-ended question formats, enabling a precise and nuanced assessment of a model's reasoning across varying levels of difficulty. Through rigorous testing of 19 state-of-the-art open-source (general-purpose and reasoning) and closed-source LMMs, we show that top-performing models remain biased on several social stereotypes, and we demonstrate that "thinking" models induce more bias in their reasoning chains. This benchmark represents a significant step toward fostering fairness in AI systems and reducing harmful biases, laying the groundwork for more equitable and socially responsible LMMs. Our dataset and evaluation code are released publicly.

BBQ-V provides a more rigorous and standardized framework for evaluating visual stereotype bias in LMMs.

Main contributions:
  • We introduce BBQ-V, a diverse open-ended benchmark featuring 14,144 non-synthetic visual samples spanning nine categories and 50 sub-categories of social biases, providing a more accurate reflection of real-world contexts.
  • BBQ-V is meticulously designed around visually grounded scenarios, explicitly disentangling visual biases from textual biases. This enables a focused and precise evaluation of visual stereotypes in LMMs.
  • We benchmark 19 state-of-the-art open- and closed-source general-purpose and reasoning LMMs, along with their scale variants, on BBQ-V. Our analysis highlights critical challenges and provides actionable insights for developing more equitable multimodal models.
  • We further compare our setup against synthetic images and closed-ended (MCQ) evaluations, highlighting distribution shift and selection bias, respectively.

BBQ-V Dataset Overview

Benchmark comparison table

Table: Comparison of various LMM evaluation benchmarks focused on stereotypical social biases. Our proposed benchmark, BBQ-V, assesses nine social bias types and is based on real images. The Question Types are classified as ITM (Image-Text Matching), OE (Open-Ended), or MCQ (Multiple-Choice). Real Images indicates whether the dataset was synthetically generated, Image Variations refers to multiple variations of a single context, and Multi-Actors indicates whether the images contain multiple people. Compared to the widely used non-synthetic VLStereoSet benchmark, BBQ-V offers 4× more visual images, 7× more evaluation questions, and 2× more bias domains.

BBQ-V comprises nine social bias categories.

Bias types table

Table: Bias Types: Examples from the nine bias categories. The source which attests the bias is reported.

Dataset Collection and Verification Pipeline

BBQ-V data curation pipeline

Figure — BBQ-V Data Curation Pipeline: Our benchmark incorporates ambiguous contexts and bias-probing questions from the BBQ [Parrish et al., 2021] dataset. The ambiguous text context is passed to a Visual Query Generator (VQG), which simplifies it into a search-friendly query to retrieve real-world images from the web. Retrieved images are filtered through a three-stage process: (1) PaddleOCR eliminates text-heavy images; (2) semantic alignment is verified using CLIP, Qwen2.5-VL, and GPT-4o-mini to ensure the image matches the simplified context; and (3) synthetic and cartoon-like images are removed using GPT-4o-mini. A Visual Information Remover (VIR) anonymizes text references to prevent explicit leakage. The processed visual content is first blurred to remove personally identifiable data (e.g., faces, watermarks) and then paired with the original bias-probing question to construct the multimodal bias evaluation benchmark.

Data Statistics

BBQ-V spans nine diverse social bias categories and 50 sub-domains. Starting from 5,194 ambiguous BBQ context-question pairs and an over-retrieve-then-filter strategy, the pipeline yields 4,497 high-quality real images and over 14.1K image-question pairs. The dataset is demographically balanced, with near gender parity (50.9% F / 49.1% M) and broad coverage of nationalities (50.3% American, 26.4% Asian, 11.8% MENA) and ethnic identities (24.5% Caucasian, 23.3% Hispanic, 18.4% African American, 5.9% Arab).

Qualitative failure cases

Figure: We present qualitative failure cases across stereotype categories in BBQ-V from proprietary, open-source, and thinking models. Rather than refusing to answer when faced with ambiguous or insufficient information, models often rely on stereotypical associations to make definitive choices. For instance, models infer household responsibility based on traditional attire, or assume that a secretary is "often female"—both reflecting bias-driven reasoning rather than grounded inference. These examples highlight how current LMMs tend to amplify or reproduce social stereotypes when interpreting vague or context-light scenarios.

Experimental results on BBQ-V

Evaluation Framework

To evaluate stereotype bias beyond the limitations of multiple-choice QA, we adopt a reasoning-based, LLM-as-a-judge framework for open-ended multimodal responses. For each sample, the model generates a free-form rationale grounded in the image and context, plus a final answer. The judge model (GPT-4.1-mini) then scores the output on five independent 1–5 dimensions: Fairness (does the outcome avoid disadvantaging any group), Stereotype (does the reasoning rely on group-based generalizations), Prior Bias (use of unstated, non-group cultural assumptions), Ambiguity Recognition (awareness of missing evidence and willingness to refuse), and Faithfulness (perceptual grounding and absence of hallucination). The overall score is the harmonic mean across categories, where higher scores indicate more fair (non-stereotypical) outputs. A human verification study over 1,000 QA pairs confirms reliability, with 0.91 observed agreement and a 0.81 weighted Cohen's κ against domain experts.

Performance of Open-, Reasoning-, and Closed-Source LMMs on BBQ-V

Below, we present category-wise fairness results for open-source, reasoning, and closed-source models on BBQ-V.

Main results table

Evaluation of open-source, thinking-mode, and proprietary LMMs on demographic fairness metrics. Higher scores indicate more fair (non-stereotypical) outputs across demographic categories. The Average column reports the harmonic mean over all categories.

Main findings

We benchmark 19 state-of-the-art LMMs on BBQ-V (16 general-purpose and three reasoning models), evaluating across model families and scales. Our analysis highlights persistent performance gaps and biases. BBQ-V consists of over 14,100 image-question pairs.

1) Overall Results. Proprietary models lead, with Gemini-2.5-Flash-Lite (80.89%) and GPT-4o (75.38%) achieving the highest overall scores, particularly in challenging categories such as Nationality, Race/Ethnicity, and Sexual Orientation. The strongest open-source models are Phi-4-Multimodal-Instruct (74.30%), Phi-3.5-Vision-Instruct (72.59%), and Qwen2.5-Omni-7B (72.00%). In contrast, models such as LLaMA-3.2-Vision-11B (43.36%), Molmo-7B (45.03%), and LLaVA-OneVision-7B (53.19%) struggle across most categories. Even the best models show uneven performance, struggling most with Physical Appearance, Age, and Disability, while performing better on Socio-Economic Status, Nationality, and Race/Ethnicity.

2) Thinking leads to more biased responses. The reasoning-oriented models—GLM-4.1V-9B-Thinking (53.58%), SophiaVL-R1 (65.51%), and Qwen3-VL-8B-Thinking (69.47%)—lag behind similarly sized non-thinking LMMs. Qualitative inspection shows their multi-step reasoning chains do not translate into more robust behavior under ambiguity; instead, they construct narratives that link visual cues to latent attributes (e.g., economic status, family aspirations, submissiveness), thereby exposing and amplifying pre-existing social priors.

3) Assessing bias across modalities. A blind-vs-vision ablation shows that adding visual input amplifies bias relative to text-only base LLMs. In the blind setting, models lack the visual context and tend to safely refuse; once images are added, scores drop sharply: Gemma-3-12B 80.06% → 63.41%, LLaMA-3.2-11B-V 62.80% → 43.36%, Qwen2.5-VL-7B 81.19% → 65.53%, and Phi-4-MM-Instruct 89.44% → 74.30%. The drop is statistically significant across all families (p < 0.001 for paired t-test and Wilcoxon).

4) Impact of model scale on stereotype biases. Larger LMMs generally exhibit improved fairness. Across InternVL-3 (8B vs. 78B), Qwen2.5-VL (7B vs. 72B), and GPT-4o (Mini vs. Full), the larger variant consistently scores higher. For example, GPT-4o improves over GPT-4o-mini on Disability Status (72.05% vs. 50.64%) and Socio-Economic Status (84.31% vs. 72.45%), and Qwen2.5-VL-72B outperforms its 7B counterpart on Race/Ethnicity (79.25% vs. 69.20%).

5) Closed-ended (MCQ) evaluation. A multiple-choice ablation following the BBQ format lifts many models into the 70–80% range, but the higher scores reflect inflated performance from forced selection between provided options. Models with moderate open-ended performance also show high MCQ scores, suggesting selection bias rather than genuine fairness, and confirming the value of open-ended evaluation.

6) Comparison with synthetic images. Synthetic images introduce substantial distribution shift. A synthetic-minus-real analysis reveals large, structured deviations—especially in Race/Ethnicity and Socio-Economic Status (e.g., differences of 19.9% and 13.6% for Gemma-3-12B and Qwen2.5-VL-7B). Synthetic generations often fail to capture the intended visual context, motivating our reliance on real images and treating synthetic experiments only as coarse trend checks.

Conclusion

We introduce BBQ-V, a benchmark for evaluating stereotype biases in Large Multimodal Models through visually grounded contexts. BBQ-V contains over 14.1k non-synthetic, open-ended VQA pairs spanning nine domains and 50 sub-domains. We evaluate 19 LMMs, including three thinking models, and analyze three major model families (InternVL-3, Qwen2.5-VL, and GPT-4o) across multiple scales, revealing substantial performance disparities. We compare real versus synthetic images, showing that synthetic data introduces distributional shifts and fails to capture real-world visual complexity, and we demonstrate the advantages of open-ended evaluation over multiple-choice formats. Our findings show that LMMs struggle most with social categories such as Physical Appearance, Age, and Disability, while performing better on Socio-Economic Status, Nationality, and Race/Ethnicity, and that higher model scaling improves overall scores. This work highlights limitations in current LMMs and identifies key directions for reducing social bias in multimodal reasoning.


BibTeX

@article{narnaware2025bbq,
  title={BBQ-V: Benchmarking visual stereotype bias in large multimodal models},
  author={Narnaware, Vishal and Vayani, Ashmal and Gupta, Rohit and Swetha, Sirnam and Shah, Mubarak},
  journal={arXiv preprint arXiv:2502.08779},
  year={2025}
}