論文の概要: When Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundary
- arxiv url: http://arxiv.org/abs/2607.11953v1
- Date: Sat, 11 Jul 2026 21:33:32 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-15 17:08:29.894407
- Title: When Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundary
- Title(参考訳): リワード状態はいつになるか : 隠れオートマトンとグループ言語境界
- Authors: Jim Allchin,
- Abstract要約: 私たちはそれをホワイトボックスの楽器で正確に答えられるようにします。
オートマトンを知ることは、最適なリターンと正確な潜伏状態という2つのことを無償で提供します。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Does a reinforcement-learning agent that earns high reward represent its task's latent state, or only a reward-correlated shortcut? The question is usually unanswerable: the "true state" is undefined. We make it exactly answerable with a white-box instrument: express the task as a hidden deterministic finite automaton (DFA), let the agent observe a symbol stream and intermittently choose the next symbol under partial control, and grant one sparse terminal reward for acceptance. Knowing the automaton gives two things for free: the optimal return (so reward becomes an interpretable normalized score) and the exact latent state at every step (so we can probe the agent's representation without ever showing it). Reward success and latent-state learning thus become separately measured quantities whose coupling is governed by three controllable axes. Optimizer strength: under weak on-policy RL the agent earns reward with the state probe at chance for every architecture, tempting the conclusion that sparse RL cannot install latent state; a pre-registered control overturns it -- PPO+GAE recovers the state, but only partially and with high seed variance. Task structure: permutation (group-language) structure is a warning sign computable from the transition function before any training, and held out on 153 capacity-controlled fresh automata it flags perception gaps at precision 0.86 (89 of 103), in one direction only. Observation informativeness: a label-free auxiliary is vacuous when observations carry no state and recovers it in proportion to how much they reveal. The payoff is a distinction reward-only evaluation cannot make: a perception gap (latent state not linearly recoverable, though representable) versus a planning gap (state recoverable but unused). High reward is thus not evidence of task understanding; whether an agent recovers latent state is predictable in advance.
- Abstract(参考訳): 高い報酬を得る強化学習エージェントは、そのタスクの潜在状態を表すのか、それとも報酬に関連のあるショートカットのみを表すのか?
問題は通常解決不可能であり、「真の状態」は定義されていない。
タスクを隠された決定論的有限オートマトン(DFA)として表現し、エージェントにシンボルストリームを観察させ、部分的な制御の下で次のシンボルを断続的に選択させ、受け入れのためのスパース端末の報酬を1つ与える。
オートマトンを知ることで、最適なリターン(報酬は解釈可能な正規化スコアになる)と、すべてのステップにおける正確な潜時状態(エージェントの表現を決して示さずに調査できる)という2つの自由な結果が得られる。
これにより、3つの制御可能な軸によって結合が制御される逆成功と潜時学習が別々に測定される。
最適化の強度: 政治上のRLが弱い場合、エージェントは全てのアーキテクチャに対して偶然に状態プローブで報酬を受け取り、スパースRLが遅延状態を設置できないという結論を誘惑する。
タスク構造: 置換(グループ言語)構造は、任意のトレーニングの前に遷移関数から計算可能な警告記号であり、153個のキャパシティ制御された新鮮なオートマトン上に保持され、精度0.86(103の89)で認識ギャップを1方向のみにフラグする。
観察情報: ラベルのない補助装置は、観測が状態を持たないときに空白であり、それらがどれだけ露出したかに比例してそれを回復する。
報酬のみの評価は区別できない: 知覚ギャップ(線形に回復可能ではないが、表現可能)と計画ギャップ(状態回復可能だが未使用)である。
したがって、高い報酬はタスク理解の証拠ではなく、エージェントが潜伏状態の回復を事前に予測可能である。
関連論文リスト
- The Need for an External Observer Formalizing the Sufficiency Gap: A Mathematical Extension of Mixture Identifiability and Contextual Grounding in Sequence Models [0.0]
我々は、決定論的テキスト構造と、保存されていない潜在状態によって支配される1つのランダム構造を持つ2元混合登録プロセスを構築した。
結果として生じるエントロピー差は通常の最適化誤差ではない。
補正信号は、その忠実度が誤解を招く状態に割り当てられたテキストのみの後方重みを超えると、テキスト履歴によって引き起こされる後続のオッズを正確に反転させる。
論文 参考訳(メタデータ) (2026-05-26T08:53:11Z) - Interpretation, Learning, and Empathy as One Constraint: A Residual-Adequacy Architecture with Accountable Abstention [0.0]
一つの量から限界が生じる小さな認知アーキテクチャを開発する。
解釈決定ユニット(IDU)は、コンテントベクトルをレギュレーションのファミリを通じて解釈し、ライセンスするアクションを決定する。
任意のコンテンツと固定構成に対して、ユニークな端末証人を持つ有限個の有界コストステップで停止する。
論文 参考訳(メタデータ) (2026-05-24T10:57:28Z) - Survive or Collapse: The Asymmetric Roles of Data Gating and Reward Grounding in Self-Play RL [76.45061154544568]
セルフプレイ強化学習は、言語モデルを独自の生成タスクで訓練し、人間ラベルなしでプロジェクタとソルバを共進化させる。
最近のシステムでは強い推理効果が報告されているが、崩壊と不安定性は広く観察され、理解されていない。
代わりに、自己プレイの安定性は、提案者生成タスクがトレーニングプールに入るかを判断するデータレベルゲートと、すでに認められたタスクに関するポリシーを更新する報酬信号の2つの異なるレバーによって管理されていると論じる。
論文 参考訳(メタデータ) (2026-05-21T09:19:23Z) - Pando: Do Interpretability Methods Work When Models Won't Explain Themselves? [53.07826484214082]
モデル・オーガニゼーションのベンチマークであるPandoを紹介します。
Pandoは、ラベル付きクエリ-レスポンスペアから、ホールドアウトモデル決定を予測する。
説明が忠実であれば、ブラックボックスの引用はすべてのホワイトボックスメソッドに一致するか、超える。
論文 参考訳(メタデータ) (2026-04-13T06:42:24Z) - Efficient Reinforcement Learning with Impaired Observability: Learning
to Act with Delayed and Missing State Observations [92.25604137490168]
本稿では,制御系における効率的な強化学習に関する理論的研究を紹介する。
遅延および欠落した観測条件において,RL に対して $tildemathcalO(sqrtrm poly(H) SAK)$ という形でアルゴリズムを提示し,その上限と下限をほぼ最適に設定する。
論文 参考訳(メタデータ) (2023-06-02T02:46:39Z) - Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration [97.19464604735802]
探索のための有望な技術は、訪問状態分布のエントロピーを最大化することである。
エージェントが高価値の状態を訪問することを好むような、タスク報酬を伴う教師付きセットアップで苦労する傾向があります。
本稿では,値条件のエントロピーを最大化する新しい探索手法を提案する。
論文 参考訳(メタデータ) (2023-05-31T01:09:28Z) - Learning One Representation to Optimize All Rewards [19.636676744015197]
我々は,報酬のないマルコフ決定プロセスのダイナミクスのフォワードバックワード(fb)表現を紹介する。
後尾に指定された報酬に対して、明確な準最適ポリシーを提供する。
これは任意のブラックボックス環境で制御可能なエージェントを学ぶためのステップです。
論文 参考訳(メタデータ) (2021-03-14T15:00:08Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。