論文の概要: Attribute-Based Activation Steering of LLMs for Group-Specific Explanation Generation
- arxiv url: http://arxiv.org/abs/2608.29215v2
- Date: Tue, 01 Sep 2026 07:35:44 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-02 14:14:25.942847
- Title: Attribute-Based Activation Steering of LLMs for Group-Specific Explanation Generation
- Title(参考訳): グループ特異的説明生成のためのLCMの属性ベース活性化ステアリング
- Authors: Leandra Fichtel, Janek Prange, Henning Wachsmuth,
- Abstract要約: 本稿では,まず,特定の対象グループの説明スタイルと知識の観点から,グループ固有の属性を識別する手法を提案する。
実験では、生成した説明の特異性と事実性の観点から、ステアリングの有効性を評価する。
- 参考スコア(独自算出の注目度): 12.009555264358182
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: To effectively enable people to understand new topics, explanations should be tailored to their backgrounds and abilities. Prompting alone has been shown to be insufficient for creating such explanations, and other computational methods are missing so far. Therefore, this paper investigates whether LLMs can be steered to generate explanations that are tailored to a specific group of people. To this end, we propose an approach that first identifies group-specific attributes in terms of explanatory style and knowledge of a specific target group. Building on activation engineering, it then computes attribute-based steering vectors and adds them to the internal activations of an LLM during inference to enable a fine-grained steering. In our experiments, we assess the steering effectiveness in terms of specificity and factuality of the generated explanations. Additionally, we evaluate the explanations in a study with human experts from different target groups. Compared to prompting and state-of-the-art steering baselines, our approach tailors the explanations significantly better to the target group while maintaining the best specificity-factuality balance.
- Abstract(参考訳): 人々が新しいトピックを効果的に理解できるように、説明は彼らのバックグラウンドと能力に合わせて調整されるべきである。
プロンプティングだけではそのような説明を作成するには不十分であることが示されており、他の計算方法が欠落している。
そこで本稿では, LLM が特定のグループに適合した説明を作成できるかどうかを考察する。
そこで本研究では,まず,特定の対象グループの説明スタイルと知識の観点から,グループ固有の属性を識別する手法を提案する。
アクティベーションエンジニアリングに基づいて、属性ベースのステアリングベクトルを計算し、推論中にLSMの内部アクティベーションに追加して、きめ細かいステアリングを可能にする。
本実験では,生成した説明の特異性と事実性の観点から,ステアリングの有効性を評価する。
また,異なる対象グループの人間専門家による研究において,その説明を評価した。
提案手法は,プロンプトベースラインや最先端ステアリングベースラインと比較して,最適特異性と実効性バランスを維持しつつ,対象グループに対して極めて良好な説明を行う。
関連論文リスト
- From latent factors to language: a user study on LLM-generated explanations for an inherently interpretable matrix-based recommender system [8.280161440212504]
大規模言語モデル(LLM)が数学的に解釈可能なレコメンデーションモデルから,効果的なユーザ向け説明を生成できるかどうかを検討する。
本研究は,5次元にわたる説明の質を評価する326人の被験者を対象に実施した。
分析の結果、全ての説明型は概ね好意的であり、戦略間の統計的差異は緩やかであることがわかった。
論文 参考訳(メタデータ) (2025-09-23T13:30:03Z) - CohEx: A Generalized Framework for Cohort Explanation [5.269665407562217]
コホートの説明は、特定のグループや事例のコホートにおける説明者の振る舞いに関する洞察を与える。
本稿では,コホートの説明を測る上でのユニークな課題と機会について論じる。
論文 参考訳(メタデータ) (2024-10-17T03:36:18Z) - Local Explanations and Self-Explanations for Assessing Faithfulness in black-box LLMs [1.03590082373586]
本稿では,局所摂動と自己説明を用いた大規模言語モデル(LLM)の忠実度を評価するための新しい課題を紹介する。
提案手法は, 従来から用いられてきた手法にインスパイアされた, より効率的な代替的説明可能性手法を提案する。
論文 参考訳(メタデータ) (2024-09-18T10:16:45Z) - Evaluating Human Alignment and Model Faithfulness of LLM Rationale [66.75309523854476]
大規模言語モデル(LLM)が,その世代を理論的にどのように説明するかを考察する。
提案手法は帰属に基づく説明よりも「偽り」が少ないことを示す。
論文 参考訳(メタデータ) (2024-06-28T20:06:30Z) - Learning to Generate Explainable Stock Predictions using Self-Reflective
Large Language Models [54.21695754082441]
説明可能なストック予測を生成するために,LLM(Large Language Models)を教えるフレームワークを提案する。
反射剤は自己推論によって過去の株価の動きを説明する方法を学ぶ一方、PPOトレーナーは最も可能性の高い説明を生成するためにモデルを訓練する。
我々のフレームワークは従来のディープラーニング法とLLM法の両方を予測精度とマシューズ相関係数で上回ることができる。
論文 参考訳(メタデータ) (2024-02-06T03:18:58Z) - Complementary Explanations for Effective In-Context Learning [77.83124315634386]
大規模言語モデル (LLM) は、説明のインプロンプトから学習する際、顕著な能力を示した。
この研究は、文脈内学習に説明が使用されるメカニズムをよりよく理解することを目的としている。
論文 参考訳(メタデータ) (2022-11-25T04:40:47Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。