論文の概要: From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents
- arxiv url: http://arxiv.org/abs/2604.19775v1
- Date: Fri, 27 Mar 2026 22:29:01 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-05-04 02:32:14.067269
- Title: From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents
- Title(参考訳): 行動から理解へ:LLMエージェントにおける時間的概念のコンフォーマル解釈可能性
- Authors: Trilok Padhi, Ramneet Kaur, Krishiv Agarwal, Adam D. Cobb, Daniel Elenius, Manoj Acharya, Colin Samplawski, Alexander M. Berenbeim, Nathaniel D. Bastian, Susmit Jha, Anirban Roy,
- Abstract要約: 大規模言語モデル(LLM)は、対話的な環境で推論、計画、行動が可能な自律エージェントとして、ますます多くデプロイされている。
本稿では,LLMエージェントにおける概念の時間的進化をステップワイドな共形レンズで解釈する枠組みを提案する。
- 参考スコア(独自算出の注目度): 45.1959078540948
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents capable of reasoning, planning, and acting within interactive environments. Despite their growing capability to perform multi-step reasoning and decision-making tasks, internal mechanisms guiding their sequential behavior remain opaque. This paper presents a framework for interpreting the temporal evolution of concepts in LLM agents through a step-wise conformal lens. We introduce the conformal interpretability framework for temporal tasks, which combines step-wise reward modeling with conformal prediction to statistically label model's internal representation at each step as successful or failing. Linear probes are then trained on these representations to identify directions of temporal concepts - latent directions in the model's activation space that correspond to consistent notions of success, failure or reasoning drift. Experimental results on two simulated interactive environments, namely ScienceWorld and AlfWorld, demonstrate that these temporal concepts are linearly separable, revealing interpretable structures aligned with task success. We further show preliminary results on improving an LLM agent's performance by leveraging the proposed framework for steering the identified successful directions inside the model. The proposed approach, thus, offers a principled method for early failure detection as well as intervention in LLM-based agents, paving the path towards trustworthy autonomous language models in complex interactive settings.
- Abstract(参考訳): 大規模言語モデル(LLM)は、対話的な環境で推論、計画、行動が可能な自律エージェントとして、ますます多くデプロイされている。
多段階の推論と意思決定のタスクを実行する能力の増大にもかかわらず、そのシーケンシャルな振る舞いを導く内部メカニズムはいまだ不透明である。
本稿では,LLMエージェントにおける概念の時間的進化をステップワイドな共形レンズで解釈する枠組みを提案する。
本稿では,各ステップにおけるモデルの内部表現を成功または失敗として統計的にラベル付けし,段階的報酬モデルと共形予測を組み合わせ,時間的タスクに対する共形解釈可能性フレームワークを提案する。
線形プローブはこれらの表現に基づいて訓練され、時間的概念の方向 - モデルのアクティベーション空間における、成功、失敗、推論ドリフトという一貫した概念に対応する潜在方向 - を識別する。
シミュレーションされた2つの対話環境、すなわちScienceWorldとAlfWorldの実験結果は、これらの時間的概念が線形に分離可能であることを示した。
さらに, LLMエージェントの性能向上のための予備的な結果を示す。
提案手法は, 早期故障検出の原理的手法と, LLMをベースとしたエージェントの介入により, 複雑な対話的環境下での信頼性の高い自律言語モデルへの道を拓く。
関連論文リスト
- LLMs Reading the Rhythms of Daily Life: Aligned Understanding for Behavior Prediction and Generation [53.62804271492357]
大きな言語モデル(LLM)は、その意味的豊かさ、強い解釈可能性、生成能力により、有望な方向性を提供する。
本稿では,LLMを構造化カリキュラム学習プロセスを通じて人間行動モデリングに統合する,行動理解アライメント(BUA)を提案する。
BUAは、事前訓練された行動モデルからのシーケンス埋め込みをアライメントアンカーとして採用し、3段階のカリキュラムを通じてLLMをガイドし、マルチラウンドの対話設定では予測と生成機能を導入している。
論文 参考訳(メタデータ) (2026-04-26T07:34:37Z) - Toward Formalizing LLM-Based Agent Designs through Structural Context Modeling and Semantic Dynamics Analysis [13.919694566467053]
この断片化は、LLMエージェントの特性と比較を可能にする分析可能な自己整合形式モデルが存在しないことに起因すると我々は主張する。
このギャップに対処するために、文脈構造の観点からLLMエージェントを解析・比較するための形式モデルであるtexttt Structure Context Model を提案する。
モンキー・バナナ問題の動的変種に対する完全な枠組みの有効性を実証し,本手法を用いて開発したエージェントが成功率を最大32ポイント向上することを示した。
論文 参考訳(メタデータ) (2026-02-09T05:15:11Z) - Emergent Structured Representations Support Flexible In-Context Inference in Large Language Models [77.98801218316505]
大型言語モデル(LLM)は、人間のような推論を示唆する創発的な行動を示す。
テキスト内概念推論におけるLLMの内部処理について検討する。
論文 参考訳(メタデータ) (2026-02-08T03:14:39Z) - The Landscape of Agentic Reinforcement Learning for LLMs: A Survey [103.32591749156416]
エージェント強化学習(Agentic RL)の出現は、大規模言語モデル(LLM RL)に適用された従来の強化学習からパラダイムシフトを示している。
本研究では, LLM-RLの縮退した単段階マルコフ決定過程(MDPs)と, エージェントRLを定義する部分可観測マルコフ決定過程(POMDPs)とを対比することにより, この概念シフトを定式化する。
論文 参考訳(メタデータ) (2025-09-02T17:46:26Z) - Concept Layers: Enhancing Interpretability and Intervenability via LLM Conceptualization [2.163881720692685]
本稿では,概念層をアーキテクチャに組み込むことにより,解釈可能性とインターベンタビリティを既存モデルに組み込む新しい手法を提案する。
我々のアプローチは、モデルの内部ベクトル表現を、再構成してモデルにフィードバックする前に、概念的で説明可能なベクトル空間に投影する。
複数のタスクにまたがるCLを評価し、本来のモデルの性能と合意を維持しつつ、意味のある介入を可能にしていることを示す。
論文 参考訳(メタデータ) (2025-02-19T11:10:19Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。