論文の概要: When a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models
- arxiv url: http://arxiv.org/abs/2607.14169v2
- Date: Sun, 19 Jul 2026 13:17:14 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-21 14:24:57.543926
- Title: When a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models
- Title(参考訳): 検証された世界モデルが依然として失われる時: LLM合成コードワールドモデルにおけるプレイ精度と予測精度
- Abstract要約: 大規模な言語モデルは、ゲームのルールを実行可能なコードとして合成することができる。
これらは典型的には、サンプリングされた軌道上で高い遷移精度に達すると受け入れられる。
これは計画上の適切性という間違った概念である、と我々は主張する。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Large language models can synthesize a game's rules as executable code - a Code World Model (CWM) - which a classical planner then searches over. Such models are typically accepted when they reach high transition accuracy on sampled trajectories. We argue this is the wrong notion of adequacy for planning. We show four things. (1) An LLM-synthesized CWM can pass a sampling gate at 100% transition accuracy and be $\geq 98\%$ state-accurate on the planner's own search distribution, yet lose systematically at play, because the $<1\%$ it gets wrong is exactly the pivotal dynamics; the play cost of the omitted rule is $0.091$ (seed-clustered 95% CI $[0.065,0.117]$, $n=4800$). We call this the verified-vs-correct gap, and confirm it end-to-end through the synthesis pipeline. (2) The harm follows a quantitative law, $\mathrm{danger}=\mathrm{play\_cost}\times(1-\mathrm{rarity})^N$, whose $(1-\mathrm{rarity})^N$ gate-miss factor is proven exact and whose play cost is empirically bounded. (3) The failure is not repaired by more data: LLM synthesis behaves as rule translation, not rule inference, and did not infer the omitted rule across models (GPT-5.x) and data regimes (including DAgger and targeted examples). (4) The same mechanism recurs on the belief-inference function of imperfect-information CWMs: we prove a coverage bound (a size-$N$ gate is identifying when $N\gtrsim b^{d_{\max}}$), explaining why shallow games such as Kuhn poker show no gap, and hand-construct Beacon, a verified-but-wrong inference function that passes the gate yet loses every game. These results suggest adequacy for planning-oriented world models should be measured on the search distribution or by play directly, not by prediction accuracy on sampled transitions.
- Abstract(参考訳): 大規模な言語モデルは、ゲームのルールを実行可能なコード(コードワールドモデル(CWM))として合成し、古典的なプランナーが検索する。
このようなモデルは通常、サンプリングされた軌道上で高い遷移精度に達すると受け入れられる。
これは計画上の適切性という間違った概念である、と我々は主張する。
私たちは4つのものを見せます。
1) LLM を合成した CWM は、サンプリングゲートを100%遷移精度で通過させ、プランナー自身の検索ディストリビューションで$\geq 98\%$を状態精度で行うことができるが、$<1\%$が間違っているのは、正確には重要なダイナミクスであるため、体系的に失われる。
これを検証-vs-正しいギャップと呼び、合成パイプラインを通じてエンドツーエンドで確認します。
2) 害は定量的な法則に従う: $\mathrm{danger}=\mathrm{play\_cost}\times(1-\mathrm{rarity})^N$, which $(1-\mathrm{rarity})^N$ gate-miss factor is confirmed exact and that play cost is empirally bounded。
LLM合成は規則翻訳として振る舞うが、規則推論ではなく、モデル(GPT-5.x)とデータ構造(DAggerや対象とする例を含む)で省略された規則を推測しなかった。
(4) 不完全情報CWMの信念推論関数に同じ機構が再帰する: カバーバウンド(サイズ-N$ gate is identified when $N\gtrsim b^{d_{\max}}$)を証明する。
これらの結果は,サンプル遷移の予測精度ではなく,探索分布やプレーによって,計画指向の世界モデルの有効性を測るべきであることを示唆している。
関連論文リスト
- Optimal Rates for Agentic Networked Information Aggregation [14.181355030230309]
ネットワーク学習モデルにおける情報集約について検討する。
エージェントはDAGに座り、それぞれが特徴とその両親の予測のサブセットしか見ず、線形予測器に適合し、予測のみを前方に通過する。
我々は,Bateni et alのロジットパスモデルにおいて,ロジスティックな分類に最適であることを示す。
論文 参考訳(メタデータ) (2026-09-04T16:12:16Z) - Reading the Room: Implicit Confusion Encoding in Recurrent World Model States [0.0]
我々は、予測エラーのトラックの混乱を軽減するために、繰り返し発生する隠れ状態$h_t$を示す。
これは、新しい入力をフラグするアンサンブルの不一致と、現在悪い予測をフラグするリコンストラクションエラーとを機能的に区別している。
論文 参考訳(メタデータ) (2026-08-21T19:32:38Z) - A Mathematical Theory of Value: a synthesis on goal-directed agency under resource constraints [6.057587531186626]
目的指向エージェントが生成し、破壊し、交換する価値は、情報と同じカテゴリの法的構造量であることを示す。
価格がフレームに依存していない間、価値はフレーム相対的であり、そのリソースをプールし、その知覚を融合する艦隊が天井を継承する。
論文 参考訳(メタデータ) (2026-06-10T16:11:04Z) - The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations [50.43168858368539]
大規模言語モデルは自信を持って時代遅れの回答を生成し、既存の方法では検出できない。
これは工学的な失敗ではなく構造的な失敗であり、時間的ドリフトは、幾何的に残留流の方向として、正確性と不確実性の両方に符号化される。
論文 参考訳(メタデータ) (2026-05-09T22:27:31Z) - Scale-Invariant Regret Matching and Online Learning with Optimal Convergence: Bridging Theory and Practice in Zero-Sum Games [60.871651115241406]
ゼロサムゲームにおける理論と実践の間、何十年にもわたってかなりのシャズムが一階法によって浸食されてきた。
我々は、IREG-PRM$+$と呼ぶPRM$+$の新しいスケール不変かつパラメータフリーな変種を提案する。
ベンチマークゲームでは, PRM$+$と同等でありながら, 最適収束保証を$T-1/2$, $T-1$とする。
論文 参考訳(メタデータ) (2025-10-06T00:33:20Z) - Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier: Autoregressive and Imitation Learning under Misspecification [50.717692060500696]
対数損失を伴う次のトーケン予測は自己回帰シーケンスモデリングの基盤となる。
次トーケン予測は、適度な誤差増幅を表す$C=tilde O(H)$を達成するために堅牢にすることができる。
C=e(log H)1-Omega(1)$。
論文 参考訳(メタデータ) (2025-02-18T02:52:00Z) - Experts Don't Cheat: Learning What You Don't Know By Predicting Pairs [35.92045337126979]
モデルに$p(Y|X)$を近似させる戦略を提案し、$widehatp_theta(Y|X)$と$p(Y|X)$の間の残りのギャップを推定する。
提案手法では,曖昧な画像分類,(合成)言語モデリング,部分観測可能なナビゲーションタスクなどにおいて,モデルがどの程度の知識を持っていないかを正確に推定する。
論文 参考訳(メタデータ) (2024-02-13T19:01:45Z) - Minimax-Optimal Multi-Agent RL in Zero-Sum Markov Games With a
Generative Model [50.38446482252857]
2人プレイのゼロサムマルコフゲームは多エージェント強化学習においておそらく最も基本的な設定である。
我々は,$$ widetildeObiggを用いて,$varepsilon$-approximate Markov NEポリシーを学習する学習アルゴリズムを開発した。
我々は、分散型量の役割を明確にするFTRLに対する洗練された後悔境界を導出する。
論文 参考訳(メタデータ) (2022-08-22T17:24:55Z) - Topology-aware Generalization of Decentralized SGD [91.59494285490784]
D-SGDの一般化性はスペクトルギャップと正の相関関係を示す。
我々の知る限り、これはD-SGDの一般化に関する最初の研究である。
論文 参考訳(メタデータ) (2022-06-25T16:03:48Z) - Linear Contextual Bandits with Adversarial Corruptions [91.38793800392108]
本稿では,敵対的腐敗の存在下での線形文脈的包帯問題について検討する。
逆汚染レベルに適応する分散認識アルゴリズムをC$で提案する。
論文 参考訳(メタデータ) (2021-10-25T02:53:24Z) - What Happens after SGD Reaches Zero Loss? --A Mathematical Framework [35.31946061894308]
SGD(Gradient Descent)の暗黙のバイアスを理解することは、ディープラーニングにおける重要な課題の1つである。
本稿では、Katzenberger (1991) のアイデアを適応させることにより、そのような分析の一般的な枠組みを提供する。
1) a global analysis of the implicit bias for $eta-2$ steps, not to the local analysis of Blanc et al. (2020) that is only for $eta-1.6$ steps and (2) allowing any noise covariance。
論文 参考訳(メタデータ) (2021-10-13T17:50:46Z) - Superdeterministic hidden-variables models II: conspiracy [0.0]
我々は、量子力学の超決定論的モデルが数学的に明確に定義された意味で相補的であることを証明した。
非平衡を使わずに超決定論的陰謀を定量化する方法を示す。
非局所モデルと後方モデルの両方のアプローチにより非補間的であることが判明した。
論文 参考訳(メタデータ) (2020-03-27T01:01:51Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。