論文の概要: WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models
- arxiv url: http://arxiv.org/abs/2608.04964v1
- Date: Wed, 05 Aug 2026 15:34:47 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-06 14:48:43.978941
- Title: WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models
- Title(参考訳): WorldCycle: 長距離ビデオワールドモデルのための自己検証型強化学習
- Authors: Bohai Gu, Yueyang Yuan, Taiyi Wu, Dazhao Du, Jian Liu, Xiaoyi Pang, Jie Zhang, Xiaocheng Lu, Haobin Zhong, Xiaotong Zhao, Alan Zhao, Song Guo,
- Abstract要約: クローズドアクションサイクルと通常のアクションシーケンスからの繰り返し実行を構成する自己検証可能なRLフレームワークであるWorldCycleを紹介する。
WorldCycleは、状態復帰ドリフトを最大44%削減し、ベースモデル上での複合動作精度をほぼ4倍に向上させる。
- 参考スコア(独自算出の注目度): 28.390938989852625
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compounding errors. Post-training methods such as reinforcement learning (RL) can improve these models, but they hit a verification bottleneck: for arbitrary action sequences, no ground-truth future state exists to measure long-term drift. Our key insight is that reversible action cycles make this verification possible: a sequence composed with its inverse must analytically return to the initial state, yielding annotation-free supervision on long-horizon correctness. Building on this, we introduce WorldCycle, a self-verifiable RL framework that constructs closed action cycles and their repeated executions from ordinary action sequences, and optimizes two complementary rewards: a spatial closure reward enforcing symmetry between mirrored forward and reverse segments, and a temporal consistency reward aligning states across repeated cycle executions. These rewards force the model to learn actions as consistent state operators rather than memorized temporal patterns, and extend naturally to out-of-distribution composite action cycles that the base model handles poorly. We further release CycleBench, a diagnostic benchmark for state-returning ability under complex action structures. WorldCycle reduces state returning drift by up to 44% and lifts composite-action accuracy nearly 4x over the base model, providing a vital foundation for physically grounded world models.
- Abstract(参考訳): インタラクティブなビデオワールドモデルは、長期の計画と探索には不可欠だが、それらは複雑なエラーに悩まされている。
強化学習(RL)のようなポストトレーニング手法はこれらのモデルを改善することができるが、それらは検証のボトルネックにぶつかっている。
我々の重要な洞察は、可逆的な行動サイクルによってこの検証が可能となり、逆数からなるシーケンスは解析的に初期状態に戻らなければならない。
これに基づいて, 閉じた動作サイクルと通常の動作シーケンスから繰り返し実行される実行を構成する自己検証可能なRLフレームワークであるWorldCycleを導入し, 2つの相補的な報酬を最適化する。
これらの報酬は、モデルを記憶された時間パターンではなく、一貫した状態演算子として学習させ、ベースモデルがうまく扱えない分配外の複合作用サイクルに自然に拡張させる。
さらに、複雑なアクション構造下での状態回復能力の診断ベンチマークであるCycleBenchをリリースする。
WorldCycleは、状態復帰のドリフトを最大44%削減し、ベースモデルの4倍近い複合動作精度を引き上げ、物理的に接地された世界モデルに不可欠な基盤を提供する。
関連論文リスト
- Cycle-World: Mitigating Error Accumulation in Long-term Video World Models via Reverse-Prediction Cycle Consistency [123.12549538776109]
Cycle-Worldは、安定的で時間的に一貫した長ビデオ生成のために設計されたフレームワークである。
提案手法は,トレーニング段階と推論段階の両方に厳密な時間的可逆性を付与することにより,誤差の漂流に対処する。
VBenchベンチマークの実験では、Cycle-Worldの2相相相のシナジーがエラードリフトを著しく緩和することを示した。
論文 参考訳(メタデータ) (2026-07-13T17:27:13Z) - AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing [45.6226084109777]
本稿では,Diffusion Transformer (DiT) アーキテクチャ上に構築されたAHA-WAM(Asynchronous Horizon-Adaptive World-Action Model)を提案する。
AHA-WAMはロボットデータの事前学習なしに最先端のパフォーマンスを達成し、RoboTwinで平均92.80%、実世界の4つのタスクで78.3%の成功を達成した。
論文 参考訳(メタデータ) (2026-06-08T17:55:18Z) - Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation [25.677744104220853]
ビデオフレームは、特定の動作セマンティクスに固執しながら、所定のエンドポイント間で現実的な中間フレームを合成することを目的としている。
本稿では,前向きと後向きの軌跡の対称性を強制する新しい双方向フレームワークを提案する。
本手法は,37フレームと73フレームの両方のタスクにおいて,画像品質,運動の滑らかさ,動的制御における最先端性能を実現する。
論文 参考訳(メタデータ) (2026-04-02T06:58:46Z) - SPIRAL: A Closed-Loop Framework for Self-Improving Action World Models via Reflective Planning Agents [135.00390535239129]
本稿では,自己改善型計画および反復的行動世界モデリングフレームワークであるSPIRALを紹介する。
SPIRALはActWMをクローズドループシンク-アクト-リフレクションプロセスとして定式化し、そこで生成は明示的な計画とフィードバックの下で段階的に進行する。
複数のTI2Vバックボーンに対する実験は、ActWM-Benchとメインストリームのビデオ生成ベンチマークで一貫した利得を示している。
論文 参考訳(メタデータ) (2026-03-09T14:00:36Z) - Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning [61.380634253724594]
次トーケン予測に基づく大規模自己回帰モデルの構築と強化学習(RL)による微調整
自己回帰モデルの内部表現を動作させ,探索することにより,この問題を克服できることを示す。
論文 参考訳(メタデータ) (2025-12-23T18:51:50Z) - GrndCtrl: Grounding World Models via Self-Supervised Reward Alignment [16.343768407636322]
本稿では,自己指導型ポストトレーニングフレームワークであるReinforcement Learning with World Grounding(RLWG)を紹介する。
このフレームワークをGrndCtrlでインスタンス化する。GrndCtrlは、グループ相対ポリシー最適化(GRPO)に基づく報酬整合型適応手法で、安定な軌道の維持、一貫した幾何、エンボディナビゲーションのための信頼性のあるロールアウトを行う世界モデルを生成する。
論文 参考訳(メタデータ) (2025-12-01T18:03:29Z) - Bench2Drive-R: Turning Real World Data into Reactive Closed-Loop Autonomous Driving Benchmark by Generative Model [63.336123527432136]
我々は,リアクティブ閉ループ評価を可能にする生成フレームワークであるBench2Drive-Rを紹介する。
既存の自動運転用ビデオ生成モデルとは異なり、提案された設計はインタラクティブなシミュレーションに適したものである。
我々は、Bench2Drive-Rの生成品質を既存の生成モデルと比較し、最先端の性能を達成する。
論文 参考訳(メタデータ) (2024-12-11T06:35:18Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。