論文の概要: Adaptive Compute in Latent World Models: When Depth Helps, Hurts, or Doesn't Matter
- arxiv url: http://arxiv.org/abs/2607.10203v2
- Date: Tue, 14 Jul 2026 14:26:03 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-15 12:51:44.544943
- Title: Adaptive Compute in Latent World Models: When Depth Helps, Hurts, or Doesn't Matter
- Title(参考訳): 潜在世界のモデルにおける適応型計算:いつ深さが役に立つか、ハートか、重要でないか
- Abstract要約: 自己回帰的なロールアウトでは、第一の仮定は構成を生き残るために深度毎の精度を必要とする。
6/9タスクの深さはロールアウトに役立つ(本来は$$4.7times$)。
頑健な逆転はダイナミクスの特性ではなく、トレーニングによって生成される。
- 参考スコア(独自算出の注目度): 0.0
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Adaptive-compute world models -- early-exit or mixture-of-depths predictors that spend variable depth per step -- assume depth buys better predictions and can be routed adaptively. In autoregressive rollouts, the first assumption requires depth's per-step precision to survive composition. We test this with a pre-registered instrument, the shallow penalty $ρ=\mathrm{err}(\text{shallowest-exit rollout})/\mathrm{err}(\text{full-depth rollout})$, across nine DeepMind Control tasks under matched single-step ($K=1$) and multi-step ($K=4$) training, three seeds each. We find three regimes: on 6/9 tasks depth helps rollouts (intrinsic, $ρ$ up to $4.7\times$), on 2/9 the shallow exits beat the full stack (inversion, $ρ$ down to $0.85\times$), and one is flat. The robust inversion (cheetah) is not a property of the dynamics but is created by training: an ablation supervising early exits only at the first rollout step erases it ($ρ: 0.87\to1.18$, $n=8$, $Δ=+0.31$), while an intrinsic-tradeoff task is unaffected -- a double dissociation we call the routability catch-22, since the supervision that makes exits routable is what trains them to out-roll the full stack. The regime is partly predictable a priori: observation/action dimensionality and one-step model error correlate with $ρ$ at $|\text{Spearman}|\approx0.75$ ($n=9$). Inside a CEM planner, $ρ$'s sign predicts whether planning benefits from depth, most sharply on the inversion task, where shallow planning beats deep. Finally, three cautions: a task's regime depends on the metric space, the rollout horizon, and the encoder. All thresholds and gates were fixed before the compute campaign, including a pre-registered negative for the hypothesis that motivated the study.
- Abstract(参考訳): アダプティブ・コンピュテート(Adaptive-Compute)な世界モデル — ステップ毎に様々な深さを費やす、早期終了または混合深度予測モデル — は、より優れた予測を買い、適応的にルーティングできると仮定する。
自己回帰的なロールアウトでは、第一の仮定は構成を生き残るために深度毎の精度を必要とする。
我々は、これを事前登録された楽器、浅いペナルティ $ρ=\mathrm{err}(\text{shallowest-exit rollout})/\mathrm{err}(\text{full-depth rollout})$でテストする。
6/9タスクの深さはロールアウトに役立ち(本質的には$ρ$から$4.7\times$)、2/9では浅い出口がフルスタック(反転、$ρ$から$0.85\times$)を上回り、1つはフラットである。
087\to1.18$, $n=8$, $Δ=+0.31$)に対して、本質的なトレードオフタスクは影響を受けない -- 二重解離は、ルタビリティ・キャッチ-22(routability catch-22)と呼ばれる。
観察・行動次元とワンステップモデル誤差は$ρ$ at $|\text{Spearman}|\approx0.75$$$n=9$と相関する。
CEMのプランナーの中で、$ρ$'sのサインは、計画が深さから利益を得るかどうかを予測する。
最後に、3つの注意: タスクの体制はメートル法空間、ロールアウト水平線、エンコーダに依存する。
計算キャンペーンの前にすべての閾値とゲートが固定され、研究を動機づけた仮説に対して、事前登録された負の値が含まれていた。
関連論文リスト
- Exact equivariance, kept through training, buys zero-shot generalisation across the symmetry group [6.230579198456525]
等変エンコーダ$E$と等変予測器$f$から構築された潜在世界モデルは、トレーニング損失の証明可能な対称性を継承する。
このエンドツーエンドをラップトップスケール(CPU/MPS、完全シード)で検証する。
論文 参考訳(メタデータ) (2026-06-02T01:20:24Z) - Nonstationary Generalized Linear Bandits with Discounted Online Mirror Descent [39.805192541498634]
本研究では,非定常線形計算(GLBs)について検討し,期待される報酬を未知の時間変化パラメータを持つ非線形リンク関数を用いてモデル化する。
本稿では,パラメータ推定に割引オンラインミラー降下(DOMD)を利用する非定常GLBに対する新しいアルゴリズムを提案する。
論文 参考訳(メタデータ) (2026-05-25T08:40:32Z) - Dynamic Mode Decomposition along Depth in Vision Transformers [2.899294572150795]
我々は,ViTの深さがほぼ自明な線形力学を実装しているかどうかを問う。
我々は、動的モード分解(DMD)を用いてこれをテストし、選択された連続した隠れ状態ペアからK$に適合する。
予め訓練した4種類のDINO ViTについて, 安定適合に必要な正則化, ランク, 校正予算について検討した。
論文 参考訳(メタデータ) (2026-05-08T10:33:03Z) - Arithmetic-Mean $μ$P for Modern Architectures: A Unified Learning-Rate Scale for CNNs and ResNets [9.94514344279733]
Arithmetic-Mean $mu$P は個々の層ではなく、ネットワーク全体の平均1ステップのプレアクティベーション第2モーメントを一定スケールに制限する。
1次元および2次元の畳み込みネットワークの場合、最大更新学習率は$etastar(L)propto L-3/2$; を満足する。
論文 参考訳(メタデータ) (2025-10-05T19:22:50Z) - Reinforcement Learning from Adversarial Preferences in Tabular MDPs [62.73758165845971]
我々は,敵対的嗜好を持つエピソードマルコフ決定プロセス(MDP)の新たな枠組みを導入する。
PbMDP では、標準的なエピソード MDP とは異なり、学習者は2つの候補アーム間の好みを観察する。
我々は、既知遷移の下で、T2/3$という残差境界を達成するアルゴリズムを開発する。
論文 参考訳(メタデータ) (2025-07-15T20:19:32Z) - From Continual Learning to SGD and Back: Better Rates for Continual Linear Models [50.11453013647086]
以前見られたタスクの損失を、$k$の繰り返しの後、忘れること、すなわち、分析する。
実現可能な最小二乗の設定において、新しい最上界を創出する。
我々は、タスクを繰り返しないランダム化だけで、十分に長いタスクシーケンスで破滅的な事態を防げることを初めて証明した。
論文 参考訳(メタデータ) (2025-04-06T18:39:45Z) - Position-Aware Depth Decay Decoding ($D^3$): Boosting Large Language Model Inference Efficiency [26.173523821684306]
トークン配置対応層スキップフレームワークを提案し,性能を維持しつつ1.5倍の演算を効率よく節約する。
7 sim 70$のパラメータを持つ大規模言語モデルの実験では、D3$は完全な推論パイプラインと比較して平均1.5倍のスピードアップを達成することができる。
論文 参考訳(メタデータ) (2025-03-11T15:15:54Z) - Scalable 3D Registration via Truncated Entry-wise Absolute Residuals [65.04922801371363]
3ドルの登録アプローチでは、1000万ドル(107ドル)以上のポイントペアを、99%以上のランダムなアウトレイアで処理することができる。
我々はこの手法をTEARと呼び、Trncated Entry-wise Absolute Residualsを演算するoutlier-robust損失を最小限にする。
論文 参考訳(メタデータ) (2024-04-01T04:43:39Z) - Variance-Aware Confidence Set: Variance-Dependent Bound for Linear
Bandits and Horizon-Free Bound for Linear Mixture MDP [76.94328400919836]
線形バンドイットと線形混合決定プロセス(mdp)に対する分散認識信頼セットの構築方法を示す。
線形バンドイットに対しては、$d を特徴次元とする$widetildeo(mathrmpoly(d)sqrt1 + sum_i=1ksigma_i2) が成り立つ。
線形混合 MDP に対し、$widetildeO(mathrmpoly(d)sqrtK)$ regret bound を得る。
論文 参考訳(メタデータ) (2021-01-29T18:57:52Z) - Naive Exploration is Optimal for Online LQR [49.681825576239355]
最適後悔尺度は$widetildeTheta(sqrtd_mathbfu2 d_mathbfx T)$で、$T$は時間ステップの数、$d_mathbfu$は入力空間の次元、$d_mathbfx$はシステム状態の次元である。
我々の下界は、かつての$mathrmpoly(logT)$-regretアルゴリズムの可能性を排除する。
論文 参考訳(メタデータ) (2020-01-27T03:44:54Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。