論文の概要: LWDrive: Layer-Wise World-Model-Guided Vision-Language Model Planning for Autonomous Driving
- arxiv url: http://arxiv.org/abs/2606.29879v2
- Date: Tue, 30 Jun 2026 16:50:58 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-01 13:50:27.73088
- Title: LWDrive: Layer-Wise World-Model-Guided Vision-Language Model Planning for Autonomous Driving
- Title(参考訳): LWDrive: 自律運転のためのレイヤワイズ世界モデルガイド型ビジョンランゲージモデル計画
- Abstract要約: We developed the Layer-Wise World-Model-Guided Driving framework (LWDrive)
LWDriveは、階層的なワールドモデルガイダンスを通じて粗い軌跡を洗練するVLM計画フレームワークである。
実験の結果、LWDriveはNAVSIMベンチマークで92.0、NAVSIM-v2で89.6を記録した。
- 参考スコア(独自算出の注目度): 13.964957969409655
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD) planning. However, trajectories directly generated by VLMs often encode only coarse driving intentions and remain insufficient for geometrically accurate, future-aware, and multi-view-grounded planning. To address these limitations, we develop the Layer-Wise World-Model-Guided Driving framework (LWDrive). LWDrive is a VLM planning framework that refines coarse trajectories through layer-wise world-model guidance. Instead of treating the VLM output as the final trajectory, LWDrive uses it as an intent-aware coarse plan, expands a diverse candidate space around it, and progressively refines the candidates through a Foresight Cascade Planner (FCP). Specifically, we introduce future-frame generation supervision to encourage the VLM to learn forward-looking scene representations, thereby injecting planning-relevant predictive dynamics into its internal hidden states. Built upon these world-model-supervised representations, FCP exploits VLM features across multiple layers and integrates historical temporal states, Action-Query representations, and current-frame multi-view Bird's-Eye-View (BEV) features to refine candidate trajectories in a coarse-to-fine manner. This design enables progressive correction of spatial positions and motion trends while grounding trajectory refinement with multi-view scene cues and preserving the high-level driving intention produced by the large model. Finally, a score head evaluates the refined candidates and selects the best trajectory as the final planning output. Experiments show that LWDrive achieves a score of 92.0 on the NAVSIM benchmark and 89.6 on NAVSIM-v2. Code and models will be made publicly available.
- Abstract(参考訳): VLM(Vision-Language Models)は、エンドツーエンド自動運転(E2E-AD)計画のための強力なセマンティック理解とコモンセンス推論を提供する。
しかしながら、VLMが直接生成する軌道はしばしば粗い運転意図のみを符号化し、幾何学的精度、将来性、多視点の計画には不十分である。
これらの制約に対処するため、我々はLayer-Wise World-Model-Guided Driving framework (LWDrive)を開発した。
LWDriveは、階層的なワールドモデルガイダンスを通じて粗い軌跡を洗練するVLM計画フレームワークである。
VLM出力を最終軌跡として扱う代わりに、LWDriveは意図を意識した粗い計画として使用し、その周辺に多様な候補空間を拡張し、フォレスト・カスケード・プランナー(FCP)を通じて候補を徐々に洗練させる。
具体的には、将来的なフレーム生成の監督を導入し、VLMが前方のシーン表現を学習することを奨励し、内部の隠れ状態に計画関連予測ダイナミクスを注入する。
これらの世界モデルによる表現に基づいて、FCPはVLM機能を複数のレイヤにまたがって利用し、歴史的時間的状態、アクションクエリ表現、および現在の多視点Bird's-Eye-View(BEV)機能を統合することで、候補の軌跡を粗い方法で洗練する。
この設計により、多視点シーンキューによる軌道改善を基礎とし、大型モデルによる高レベルの駆動意図を保ちながら、空間的位置と動きの傾向を漸進的に補正することができる。
最後に、スコアヘッドは、精製された候補を評価し、最終計画出力として最適な軌道を選択する。
実験の結果、LWDriveはNAVSIMベンチマークで92.0、NAVSIM-v2で89.6を記録した。
コードとモデルは公開されます。
関連論文リスト
- LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model [16.87741001074065]
VLA(Vision-Language-Action)モデルは、エンドツーエンドの自動運転のための有望なフレームワークとして登場した。
近年、世界モデリングによる濃密な視覚監視を取り入れようとする試みは、しばしばピクセルレベルの画像再構成を過度に強調している。
自律運転のための遅延視覚表現拡張VLAフレームワークであるLVDriveを提案する。
論文 参考訳(メタデータ) (2026-05-21T07:31:49Z) - LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving [60.31765454895336]
本稿では、マルチモーダル理解と生成世界モデルを組み合わせた、エンドツーエンドのクローズドループ駆動のための最初のフレームワークLMGenDriveを紹介する。
本稿では,視覚前訓練から多段階長距離運転に至るまでの3段階訓練戦略を提案し,安定性と性能の向上を図る。
論文 参考訳(メタデータ) (2026-04-09T19:13:14Z) - Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving [52.04950569530877]
我々は、将来のフレーム予測と軌道計画の密接なインターリーブを行う統合視覚言語行動モデルUni-World VLAを提案する。
提案手法は,高忠実度将来のフレーム予測を行いながら,競合する閉ループ計画性能を実現する。
論文 参考訳(メタデータ) (2026-03-28T14:39:51Z) - Bridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion Representation [66.7879424097418]
We present WorldDrive, a holistic framework that couples scene generation and real-time planning through unified vision and motion representation。
動きの表現、視覚的表現、エゴ状態の間の単純な相互作用は、高品質でマルチモーダルな軌道を生成することができる。
NAVSIM、NAVSIM-v2、nuScenesベンチマークの実験は、WorldDriveが視覚のみの手法で主要な計画性能を達成することを示した。
論文 参考訳(メタデータ) (2026-03-16T07:59:39Z) - NaviDriveVLM: Decoupling High-Level Reasoning and Motion Planning for Autonomous Driving [4.400011068855375]
本研究では,大規模ナビゲータと軽量トレーニングドライバを用いた行動生成から推論を分離するフレームワークであるNaviDriveVLMを提案する。
nuScenesベンチマークの実験では、NaviDriveVLMはエンド・ツー・エンドの動作計画において大きなVLMベースラインを上回っている。
論文 参考訳(メタデータ) (2026-03-09T02:47:44Z) - AppleVLM: End-to-end Autonomous Driving with Advanced Perception and Planning-Enhanced Vision-Language Models [11.748457186467727]
我々は、堅牢なエンドツーエンド運転のための先進的な認識と計画強化VLMモデルであるAppleVLMを提案する。
AppleVLMは、新しいビジョンエンコーダと計画戦略エンコーダを導入し、認識と意思決定を改善する。
我々は,CARLAベンチマークのクローズドループ実験において,AppleVLMを評価し,最先端の駆動性能を実現する。
論文 参考訳(メタデータ) (2026-02-04T06:37:14Z) - SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Driving [52.02379432801349]
本稿では,運転特化知識階層に関するVLMの表現学習を構築する新しいフレームワークであるSGDriveを提案する。
トレーニング済みのVLMバックボーン上に構築されたSGDriveは、人間の運転認知を反映するシーンエージェントゴール階層に、駆動理解を分解する。
論文 参考訳(メタデータ) (2026-01-09T08:55:42Z) - UniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous Driving [29.623672055601418]
本稿では,運転シーン理解,軌道計画,軌跡条件付き将来の画像生成を共同で行う,統一VLMベースの世界モデルを提案する。
Bench2Driveベンチマークの実験では、UniDrive-WMは高忠実な将来の画像を生成し、L2軌道誤差が5.9%、衝突速度が9.2%向上している。
論文 参考訳(メタデータ) (2026-01-07T23:49:52Z) - Enhancing End-to-End Autonomous Driving with Latent World Model [78.22157677787239]
本稿では,LAW(Latent World Model)を用いたエンドツーエンド運転のための自己教師型学習手法を提案する。
LAWは、現在の特徴とエゴ軌道に基づいて将来のシーン機能を予測する。
この自己監督タスクは、知覚のない、知覚に基づくフレームワークにシームレスに統合することができる。
論文 参考訳(メタデータ) (2024-06-12T17:59:21Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。