論文の概要: Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations
- arxiv url: http://arxiv.org/abs/2607.28319v1
- Date: Thu, 30 Jul 2026 14:54:33 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-31 21:37:00.614631
- Title: Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations
- Title(参考訳): フェアネス・プルーニング:GLU-MLP層におけるディファレンシャルアクティベーションによるデモグラフィックバイアスの配置
- Authors: Pere Martra, Eugenio Martínez Cámara, Alfonso Ureña López,
- Abstract要約: 本研究は,大規模言語モデル(LLM)における階層バイアスの管理と緩和を目的とした構造的介入手法であるFairness Pruningを提示する。
最小コントラストのプロンプトペアと推論時アクティベーションキャプチャを用いて、GLUアーキテクチャの階層特性を処理する際に異なる反応をするニューロンを同定する。
その結果、同定されたニューロンをゼロにすると、モデルが関連する人口統計学変数にどのように反応するかが変化することが示された。
- 参考スコア(独自算出の注目度): 0.21410799064827235
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: This work presents Fairness Pruning, a lightweight structural intervention method designed for the management and future mitigation of demographic bias in large language models (LLMs). As a foundational empirical validation of this method, this work focuses on causal bias localization. Using minimally contrastive prompt pairs and inference-time activation capture, the method identifies neurons that react differentially when processing demographic attributes in GLU architectures, evaluating the signal at the down_proj input. Empirical evaluation was conducted on models of up to 3 billion parameters (Llama-3.2 family and Salamandra-2B), combining standardized benchmark evaluation with qualitative text generation experiments. Results demonstrate that zeroing the identified neurons alters how the model responds to associated demographic variables. However, rather than producing flat mitigation, the intervention causes bidirectional bias destabilization: because BiasScore is unsigned, candidate sets mix neurons that push toward and against the stereotype, and the net effect on aggregate bias depends on which sign dominates. The intervention is extremely surgical: zeroing at most 40 neurons in Llama-3.2-1B (less than 0.031% of total MLP width) achieves a mean retention of 99.49% in reasoning and general knowledge capabilities. These findings empirically confirm that demographic bias processing and model capabilities operate on dissociable circuits, establishing the methodological foundations for transitioning from blind zeroing toward directional behavior modulation.
- Abstract(参考訳): 本研究は,大規模言語モデル(LLM)における階層バイアスの管理と緩和を目的とした,軽量な構造介入手法であるFairness Pruningを紹介する。
本手法の基礎的実証的検証として,因果バイアスの局在に着目した。
最小のコントラッシブなプロンプトペアと推論時アクティベーションキャプチャを用いて、GLUアーキテクチャの階層特性を処理する際に異なる反応をするニューロンを特定し、down_proj入力で信号を評価する。
最大30億のパラメータ(Llama-3.2 family と Salamandra-2B )のモデルで実験的な評価を行い、標準化されたベンチマーク評価と定性テキスト生成実験を組み合わせた。
その結果、同定されたニューロンをゼロにすると、モデルが関連する人口統計学変数にどのように反応するかが変化することが示された。
しかし、この介入は平らな緩和を生成するのではなく、双方向のバイアス不安定化を引き起こす: BiasScore は符号のないため、候補セットはステレオタイプを向いたり向いたりするニューロンを混合し、集合バイアスに対するネット効果はどのシグナルが支配するかに依存する。
Llama-3.2-1B の40ニューロン(総 MLP 幅の 0.031% 未満)をゼロにすると、推論と一般的な知識能力において99.49%の保持が達成される。
これらの知見は, 集団バイアス処理とモデル能力が解離可能な回路で動作していることを実証的に確認し, ブラインドゼロ化から指向性変調へ移行するための方法論の基礎を確立した。
関連論文リスト
- Unraveling Machine Behavior by Multi-Level Bias Analysis and Detection: Methodology and Application to Computer Vision [10.595890561268735]
本研究では,ニューラルネットワークにおけるバイアスの存在と伝播について,包括的多段階解析を用いて検討する。
本研究では,SpaceBias,ActivationBias,WeightBiasの3つのバイアス検出手法を提案する。
提案手法は,ネットワークアーキテクチャ自体のバイアスが,異なるレベルでどのように現れるかを示す。
論文 参考訳(メタデータ) (2026-07-08T10:18:51Z) - SL-BiLEM: Structured Learnable Behavior-in-the-Loop Epidemic Modeling for Forecasting and Policy Evaluation [42.644852156101685]
エピデミック予測は、人間の行動が病気の拡散に動的に反応する、という根本的な課題に直面している。
頑健な外挿のための正則化として物理的制約を利用するtextbfSL-BiLEM(Structured Learnable Behavior-in-the-Loop Epidemic Model)を提案する。
論文 参考訳(メタデータ) (2026-05-26T08:46:04Z) - Fair outputs, Biased Internals: Causal Potency and Asymmetry of Latent Bias in LLMs for High-Stakes Decisions [0.0]
本稿では,人種的関係の異なるアプリケーションを用いて,オープンウェイトモデルを用いた住宅ローン引受について検討する。
モデルでは, 出力レベルの偏りは見られず, モデル層全体にわたる人口動態の表現を保ち, 増幅している。
アクティベーションステアリングと新しい層間干渉により、この抑圧された情報が決定に関連があることを実証する。
論文 参考訳(メタデータ) (2026-05-12T12:14:58Z) - Fairness Testing in Retrieval-Augmented Generation: How Small Perturbations Reveal Bias in Small Language Models [0.1876920697241348]
本研究は, 3つの小言語モデル(SLM)による感情分析における公平性を評価するために, 制御された人口変動を導入した変成試験(MT)による公平性テストを実施する。
その結果, 微少な人口変動が変成関係(MRs)の3分の1に分解できることが示唆された。
これらの失敗を詳細に分析すると、一貫したバイアス階層が示され、人種的手がかりを含む摂動が違反の主な原因となっている。
論文 参考訳(メタデータ) (2025-09-30T17:42:35Z) - Mitigating Multimodal Hallucinations via Gradient-based Self-Reflection [49.26064449816502]
本研究では,テキスト・視覚バイアスと共起バイアスに対処するために,グラディエントベースのインフルエンス・アウェア制約付きデコーディング(GACD)手法を提案する。
GACDは幻覚を効果的に低減し、MLLM出力の視覚的接地を改善する。
論文 参考訳(メタデータ) (2025-09-03T08:13:52Z) - CausalDiff: Causality-Inspired Disentanglement via Diffusion Model for Adversarial Defense [61.78357530675446]
人間は、本質的な要因のみに基づいて判断するので、微妙な操作によって騙されるのは難しい。
この観察に触発されて、本質的なラベル因果因子を用いたラベル生成をモデル化し、ラベル非因果因子を組み込んでデータ生成を支援する。
逆の例では、摂動を非因果因子として識別し、ラベル因果因子のみに基づいて予測することを目的としている。
論文 参考訳(メタデータ) (2024-10-30T15:06:44Z) - Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination [54.865941973768905]
本稿では,命令追従設定における言語モデルのバイアスニューロンを除去するための,新しい実用的なバイアス緩和手法であるCRISPRを提案する。
CRISPRは自動的にバイアス出力を決定し、バイアス出力に影響を与えるニューロンを説明可能性法を用いてバイアスニューロンに分類する。
実験により,モデルのタスク性能と既存知識を損なうことなく,ゼロショット命令追従条件下でのバイアス軽減効果が示された。
論文 参考訳(メタデータ) (2023-11-16T07:16:55Z) - An Investigation of Why Overparameterization Exacerbates Spurious
Correlations [98.3066727301239]
この動作を駆動するトレーニングデータの2つの重要な特性を特定します。
モデルの"記憶"に対する帰納的バイアスが,パラメータ化の超過を損なう可能性を示す。
論文 参考訳(メタデータ) (2020-05-09T01:59:13Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。