論文の概要: Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control
- arxiv url: http://arxiv.org/abs/2607.26657v2
- Date: Mon, 03 Aug 2026 03:05:35 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-04 15:07:24.001438
- Title: Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control
- Title(参考訳): 展開:超効率的な身体制御のための予測表現への世界モデルイマジネーション
- Authors: Weili Zeng, Yitong Xing, Fulong Liu, Chengqun Yang, Antao Xiang, Feng Tian, Jingnan Gao, Jisong Cai, Xin Wang, Xiaomin Wu, Yao Mu, Xiaokang Yang, Yichao Yan,
- Abstract要約: 我々は、未来を構成する計算が再利用可能な場合、世界生成モデルは最も再利用可能であると論じる。
本稿では,この計算を現在の視覚的コンテキストと言語命令から予測された表現に変換するEnfoldを提案する。
LIBERO、RoboTwin2.0、および実ロボットタスク全体で、Enfoldは強力な制御をサポートし、アクション遅延をFast--WAMと比較して3.7Times$減らし、Enfold-Flashは10.1times$に達した。
- 参考スコア(独自算出の注目度): 41.15180918658398
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: World generative models are typically used through what they produce: a rendered future, a video-conditioned action, or latent context computed by a costly generative branch. We argue that their more reusable asset is the computation that constructs a future. As a generator transforms a corrupted future into a coherent trajectory, its intermediate states organize appearance, spatial layout, and interaction across levels of abstraction. Can this future-generative computation be internalized in a representation inferred from the present alone? We present Enfold, which transfers this computation into a representation predicted from the current visual context and language instruction. During training, multi-level states exposed as the generator processes the observed future supervise a current-only encoder. The learned representation is fed back to condition future generation and is read by task heads without allowing task gradients to reshape the encoder. At deployment, action prediction no longer executes the generator. Across LIBERO, RoboTwin2.0, and real-robot tasks, Enfold supports strong control while reducing action latency by $3.7\times$ relative to Fast--WAM, Enfold-Flash reaches $10.1\times$. Representation analyses show that it suppresses nuisance variation and preferentially captures changes that emerge over longer horizons. When the current scene is altered by human intervention, both the generated continuation and the executed actions adapt, which is inconsistent with fixed trajectory replay. These results recast a world generator as a source of predictive control representations: its future need not be materialized at every step if its internal structure can be enfolded into the present.
- Abstract(参考訳): 世界生成モデルは一般的に、それらが生み出すものを通して使われる:レンダリングされた未来、ビデオ条件のアクション、またはコストのかかる生成ブランチによって計算される潜在コンテキスト。
彼らの再利用可能な資産は未来を構成する計算であると主張する。
生成元が腐敗した未来をコヒーレントな軌道に変換すると、中間状態は外観、空間的レイアウト、抽象レベル間の相互作用を整理する。
この未来生成計算は、現在のみから推測される表現に内在化できるだろうか?
本稿では,この計算を現在の視覚的コンテキストと言語命令から予測された表現に変換するEnfoldを提案する。
トレーニング中、ジェネレータとして露出したマルチレベル状態は、観測された未来が現在のみのエンコーダを監督する。
学習された表現は、将来の状態にフィードバックされ、タスクグラデーションがエンコーダを再形成することを許さずにタスクヘッドによって読み込まれる。
デプロイ時に、アクション予測はジェネレータを実行しない。
LIBERO、RoboTwin2.0、および実ロボットタスク全体で、Enfoldは強力な制御をサポートし、アクション遅延をFast--WAMと比較して3.7\times$に減らし、Enfold-Flashは10.1\times$に到達した。
表現解析により, ニュアンス変動を抑制し, より長い地平線上に出現する変化を優先的に捉えた。
現在のシーンが人間の介入によって変更されると、生成された継続と実行されたアクションの両方が適応し、固定軌跡再生と矛盾する。
これらの結果は、世界ジェネレータを予測制御表現の源として再キャストする: 内部構造が現在の状態に展開できるならば、その未来はすべてのステップで実現される必要はない。
関連論文リスト
- MIRAGE: Mobile Agents with Implicit Reasoning and Generative World Models [53.09772037247959]
可視テキスト推論トレースから連続潜時推論表現を学習するフレームワークであるMIRAGEを紹介する。
AndroidWorldでは、MIRAGEは4Bアブレーションで監督された微調整を3~5倍の低い復号化予算と一致している。
AndroidControlでは、75%以上のトークンを生成しながらアクショングラウンドを改善する。
論文 参考訳(メタデータ) (2026-06-03T09:01:24Z) - Parallel Decoding via Hidden Transfer for Lossless Large Language Model Acceleration [54.897493351694195]
本稿では,複数連続するトークンを1つのフォワードパスで同時に復号する,新しい並列復号法,すなわちthithidden Transferを提案する。
加速度測定では,Medusa や Self-Speculative decoding など,単モデル加速技術よりも優れています。
論文 参考訳(メタデータ) (2024-04-18T09:17:06Z) - iTransformer: Inverted Transformers Are Effective for Time Series Forecasting [62.40166958002558]
iTransformerを提案する。これは、逆次元に注意とフィードフォワードのネットワークを単純に適用する。
iTransformerモデルは、挑戦的な現実世界のデータセットの最先端を実現する。
論文 参考訳(メタデータ) (2023-10-10T13:44:09Z) - Future Sight: Dynamic Story Generation with Large Pretrained Language
Models [11.23192733149335]
トランスフォーマーデコーダは、以前に生成されたテキストに対してのみ新しいテキストを生成することができる。
Future Sightはデコーダが符号化された将来のプロットイベントに参加することを可能にする。
推論中、将来のプロットイベントは人間の著者によって書かれ、ある方向に生成された物語を操縦することができる。
論文 参考訳(メタデータ) (2022-12-20T01:53:26Z) - Back to the Future: Unsupervised Backprop-based Decoding for
Counterfactual and Abductive Commonsense Reasoning [79.48769764508006]
ジェネレーティブ言語モデル(LM)は、過去の文脈のみを条件にするか、狭い範囲のテキスト入力を実行するよう訓練することができる。
我々は過去と将来の両方の文脈を柔軟に組み込むことができる新しい教師なし復号アルゴリズムであるDeLoreanを提案する。
提案手法は, 帰納的テキスト生成と反事実的ストーリーリビジョンの2つの非単調推論タスクに適用可能であることを示す。
論文 参考訳(メタデータ) (2020-10-12T17:58:43Z) - Addressing Some Limitations of Transformers with Feedback Memory [51.94640029417114]
トランスフォーマーは、フィードフォワードネットワークであるにもかかわらず、シーケンシャルな自動回帰タスクにうまく適用されている。
本稿では、過去のすべての表現を将来のすべての表現に公開する、フィードバックトランスフォーマーアーキテクチャを提案する。
言語モデリング、機械翻訳、強化学習の様々なベンチマークにおいて、表現能力の増大は、同等のトランスフォーマーよりもはるかに強力なパフォーマンスを持つ、小さくて浅いモデルを生成することができることを実証する。
論文 参考訳(メタデータ) (2020-02-21T16:37:57Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。