FuguReport

ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models

Authors Shaghayegh Kolli, Sina Emami, Moreno D'IncĂ , Pouyan Nejadi, Nicu Sebe, Massimiliano Mancini, Jana Diesner
Affiliations Munich Center for Machine Learning / Munich Data Science Institute / Technical University of Munich / University of Trento / Orreco
Categories Evaluation / Bias Evaluation / Bias persistence assessment in text-to-image models, Method / Bias Evaluation / Controlled evaluation framework design, Application / Text-to-Image Generation / Bias analysis under context shift
License CC BY 4.0

Abstract Overview

This paper introduces ContextBias, a controlled evaluation framework designed to assess whether role-linked visual associations in text-to-image models persist when prompted context shifts. The authors construct ContextBench, a benchmark encompassing 92 professions and 1,656 semantically controlled prompts across context-free, related-context, and unrelated-context conditions. Evaluating four text-to-image models across 66,240 generated images, they quantify attribute stability using Bias Intensity and the Context Consistency Score. The analysis shows that contextual shifts primarily modify scene composition and camera framing, while person-level attributes, characteristic garments, and tools remain largely stable even in semantically unrelated contexts. The attribute-extraction pipeline is further supported by quantitative validation tests and a human annotation study.

Novelty

The paper provides a controlled evaluation framework and benchmark that isolates the impact of location and activity context on role-linked visual representations, rather than assessing roles solely in isolation. It introduces metrics specifically formulated to measure distributional concentration and label-level invariance across congruent, incongruent, and reformulated contextual prompts.

Results

Placing professional roles into semantically unrelated contexts does not suppress role-linked attributes; rather, pooled attribute concentration increases by +0.047 under unrelated conditions. Demographic cues, characteristic clothing, and role-specific tools exhibit high persistence across context conditions, whereas scene attributes and camera properties show greater context-sensitivity. Invariance tests indicate that 93.3% of role-label tuples remain stable under prompt paraphrases and context substitutions, with particularly high stability for People and Object attributes.

Key Points

  1. ContextBias and ContextBench establish a controlled protocol covering 92 occupational roles, 1,656 prompts, and 66,240 images to evaluate bias persistence across context shifts.
  2. Empirical evaluation across four text-to-image models demonstrates that unrelated contexts pull representations toward shared default attributes, increasing pooled Bias Intensity rather than neutralizing role priors.
  3. Person-level demographic cues, occupational garments, and tools show high persistence across context variations and prompt reformulations, while scene elements and camera parameters show greater context sensitivity.

References

This page was created using generative AI such as GPT-5, Claude Opus 4, Gemini 3, Gemini 3.1 Flash Image, and their higher-end successor versions. No guarantee can be made regarding its contents.