FuguReport

Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based Moderation

Authors Junyu Lu, Kaiyuan Liu, Jingyi Kang, Deyi Ji, Hailong Zhang, Lanyun Zhu, Qi Zhu, Bo Xu, Liang Yang, Hongfei Lin
Affiliations University of Science and Technology of China / Tencent / Dalian University of Technology / Zhejiang University / Tongji University
Categories Evaluation / Model Safety Evaluation / Robustness against rebuttal attacks, Method / Content Moderation / Judgment boundary perturbation protocol, Application / Hate Speech Detection / Mitigating false moderation risks
License CC BY 4.0

Abstract Overview

This paper studies a post-decision attack surface in LLM-based hate speech moderation: fabricated annotator-style rebuttals that ask a model to reconsider an initially correct judgment. The authors formalize two manipulation directions—whitewashing hateful content as harmless and smearing normal content as hateful—and evaluate four rebuttal strategies, including direct contradiction, boundary perturbation, adversarial rationale, and their combination. Across multiple LLMs and two hate speech datasets (SBIC and IHC), the attacks consistently reduce moderation accuracy, with compounded effects in multi-turn interactions. The analysis also demonstrates that hard-label flips understate attack severity, as rebuttals substantially erode gold-label confidence even when predictions remain correct.

Novelty

The paper introduces a rejudge protocol tailored to human–AI moderation workflows, evaluating adversarial reviewer feedback delivered after an initial model decision rather than relying solely on input-side prompt injection. It explicitly decouples attack direction (whitewashing vs. smearing) and rebuttal mechanism (boundary perturbations vs. adversarial rationales) to reveal persistent, model-specific vulnerability asymmetries.

Results

Experiments show that annotator-style rebuttals substantially degrade classification performance across models, with the strongest attack typically being boundary+rationale for GPT-5.1, Qwen3-8B, and Gemma4-E4B, and boundary alone for Gemini-2.5. Most models suffer greater degradation under smearing attacks on normal samples, and sequential rebuttals cause cumulative decline that neutral follow-ups fail to reverse. Evaluated inference-time defenses provide only partial, class-skewed recovery and leave substantial vulnerability unmitigated.

Key Points

  1. Fabricated reviewer feedback serves as an effective post-decision attack surface in collaborative moderation, successfully reversing initially correct LLM judgments.
  2. Hard-label flip rates understate attack impact because adversarial rebuttals frequently induce significant gold-label confidence erosion even when predictions do not flip.
  3. Inference-time defenses such as prior-prepending hedge and post-rebuttal prompting provide partial, class-asymmetric mitigation rather than comprehensive robustness.

References

This page was created using generative AI such as GPT-5, Claude Opus 4, Gemini 3, Gemini 3.1 Flash Image, and their higher-end successor versions. No guarantee can be made regarding its contents.