Fugu-MT 論文翻訳(概要): Quantifying Sycophancy as Deviations from Bayesian Rationality in LLMs

論文の概要: Quantifying Sycophancy as Deviations from Bayesian Rationality in LLMs

arxiv url: http://arxiv.org/abs/2508.16846v1
Date: Sat, 23 Aug 2025 00:11:00 GMT
ステータス: 翻訳完了
システム内更新日: 2025-08-26 18:43:45.208023
Title: Quantifying Sycophancy as Deviations from Bayesian Rationality in LLMs
Title（参考訳）: LLMにおけるベイズ性からの逸脱としてのシクロファンシーの定量化
Authors: Katherine Atwell, Pedram Heydari, Anthony Sicilia, Malihe Alikhani,
Abstract要約: Sycophancy, or overly agreeable or flattering behavior, is documented issue in large language model (LLMs) ベイジアンフレームワークを用いて、ユーザの視点で提示された合理的な行動からの逸脱として、梅毒を定量化する。我々は,3つの異なるタスク,オープンソースとクローズド LLM の組み合わせ,および2つの異なる方法について,サイコファンシーを探索する手法について検討した。
参考スコア（独自算出の注目度）: 26.346357679861228
License: http://creativecommons.org/licenses/by/4.0/
Abstract: Sycophancy, or overly agreeable or flattering behavior, is a documented issue in large language models (LLMs), and is critical to understand in the context of human/AI collaboration. Prior works typically quantify sycophancy by measuring shifts in behavior or impacts on accuracy, but neither metric characterizes shifts in rationality, and accuracy measures can only be used in scenarios with a known ground truth. In this work, we utilize a Bayesian framework to quantify sycophancy as deviations from rational behavior when presented with user perspectives, thus distinguishing between rational and irrational updates based on the introduction of user perspectives. In comparison to other methods, this approach allows us to characterize excessive behavioral shifts, even for tasks that involve inherent uncertainty or do not have a ground truth. We study sycophancy for 3 different tasks, a combination of open-source and closed LLMs, and two different methods for probing sycophancy. We also experiment with multiple methods for eliciting probability judgments from LLMs. We hypothesize that probing LLMs for sycophancy will cause deviations in LLMs' predicted posteriors that will lead to increased Bayesian error. Our findings indicate that: 1) LLMs are not Bayesian rational, 2) probing for sycophancy results in significant increases to the predicted posterior in favor of the steered outcome, 3) sycophancy sometimes results in increased Bayesian error, and in a small number of cases actually decreases error, and 4) changes in Bayesian error due to sycophancy are not strongly correlated in Brier score, suggesting that studying the impact of sycophancy on ground truth alone does not fully capture errors in reasoning due to sycophancy.
Abstract（参考訳）: サイコファシー(英: Sycophancy)は、大きな言語モデル(LLM)において文書化された問題であり、人間とAIのコラボレーションの文脈において理解することが重要である。以前の研究は、行動の変化や精度への影響を計測することで、典型的にはサイコフィケーシーを定量化するが、どちらの計量も合理性の変化を特徴づけておらず、精度の計測は既知の基底真理を持つシナリオでしか利用できない。本研究では,ベイズ的枠組みを用いて,ユーザの視点を提示する際の合理的行動からの逸脱として梅毒を定量化し,ユーザ視点の導入に基づく合理的かつ不合理な更新を区別する。他の手法と比較して、本手法は、本質的な不確実性を伴うタスクや根底的な真実を持たないタスクであっても、過度な行動シフトを特徴付けることができる。我々は,3つの異なるタスク,オープンソースとクローズド LLM の組み合わせ,および2つの異なる方法について,サイコファンシーを探索する手法について検討した。また,LSMから確率判断を抽出する複数の手法についても実験を行った。我々は, LLMs を梅毒に用いた場合, LLMs の後方への偏差がベイズ誤差を増大させるという仮説を立てた。私たちの発見は以下のとおりである。 1) LLM はベイズ的合理的ではない。 2) 抗酸菌症の調査では, 後部が有意に増加し, ステアリングの結果が好まれる。 3) 梅毒症は時としてベイズ誤差を増大させ, 少数の症例では実際に誤りを減少させる。 4) サイコファンシーによるベイズ誤差の変化は, ブラースコアでは強く相関せず, サイコファンシーが地上の真実のみに与える影響を研究することは, サイコファンシーによる推論における誤りを十分に捉えていないことを示唆している。

論文の概要: Quantifying Sycophancy as Deviations from Bayesian Rationality in LLMs

関連論文リスト