ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models
Abstract Overview
This paper introduces ContextBias, a controlled evaluation framework designed to assess whether role-linked visual associations in text-to-image models persist when prompted context shifts. The authors construct ContextBench, a benchmark encompassing 92 professions and 1,656 semantically controlled prompts across context-free, related-context, and unrelated-context conditions. Evaluating four text-to-image models across 66,240 generated images, they quantify attribute stability using Bias Intensity and the Context Consistency Score. The analysis shows that contextual shifts primarily modify scene composition and camera framing, while person-level attributes, characteristic garments, and tools remain largely stable even in semantically unrelated contexts. The attribute-extraction pipeline is further supported by quantitative validation tests and a human annotation study.
Novelty
The paper provides a controlled evaluation framework and benchmark that isolates the impact of location and activity context on role-linked visual representations, rather than assessing roles solely in isolation. It introduces metrics specifically formulated to measure distributional concentration and label-level invariance across congruent, incongruent, and reformulated contextual prompts.
Results
Placing professional roles into semantically unrelated contexts does not suppress role-linked attributes; rather, pooled attribute concentration increases by +0.047 under unrelated conditions. Demographic cues, characteristic clothing, and role-specific tools exhibit high persistence across context conditions, whereas scene attributes and camera properties show greater context-sensitivity. Invariance tests indicate that 93.3% of role-label tuples remain stable under prompt paraphrases and context substitutions, with particularly high stability for People and Object attributes.
Key Points
- ContextBias and ContextBench establish a controlled protocol covering 92 occupational roles, 1,656 prompts, and 66,240 images to evaluate bias persistence across context shifts.
- Empirical evaluation across four text-to-image models demonstrates that unrelated contexts pull representations toward shared default attributes, increasing pooled Bias Intensity rather than neutralizing role priors.
- Person-level demographic cues, occupational garments, and tools show high persistence across context variations and prompt reformulations, while scene elements and camera parameters show greater context sensitivity.