論文の概要: Vision-Language-Action Autonomous Driving Agent with Language-based Memory
- arxiv url: http://arxiv.org/abs/2609.38641v1
- Date: Tue, 29 Sep 2026 23:00:01 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-01 18:57:26.820783
- Title: Vision-Language-Action Autonomous Driving Agent with Language-based Memory
- Title(参考訳): 言語記憶を用いた視覚制御型自律運転エージェント
- Abstract要約: VLA(Vision-Language-Action)ファンデーションモデルが、自動運転の一般的なソリューションの1つとして登場した。
言語ベースのメモリを備えた汎用VLA駆動エージェントAD-Memoを提案する。
メモリベースのデータセットをキュレートし、2段階のレシピでVLAをトレーニングする: Supervised Fine-Tuning (SFT) と textitDa Capo は、新しい半閉鎖ループ強化学習 (RL) アルゴリズムである。
- 参考スコア(独自算出の注目度): 63.735435235930076
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Vision-Language-Action (VLA) foundation models have recently emerged as one of the prevailing solutions for autonomous driving, as they can utilize knowledge acquired during vision-language pretraining for accurate and interpretable driving. However, VLAs can take only a limited number of frames as visual input due to the high token cost of an image, which is problematic for memory-dependent tasks such as determining the arrival order at all-way stops and long-horizon driving scene understanding. Existing solutions use latent vector memories accessed through cross-attention, which are neither interpretable nor portable. In this paper, we propose AD-Memo, a general-purpose VLA driving agent with language-based memory. The agent outputs memory as an extension of its Chain-of-Thought (CoT) to record surrounding objects critical to driving; this memory becomes part of the agent's future input. We curate memory-based datasets and train VLAs with a two-stage recipe: Supervised Fine-Tuning (SFT) and \textit{Da Capo}, a novel semi-closed-loop Reinforcement Learning (RL) algorithm which uses trajectory-level advantage for memory and step-level advantage for driving, leading to better credit assignment. Across scenarios such as all-way stops and general driving, AD-Memo improves driving quality, enables better question answering on driving scenes, and produces plug-and-play memory for other models.
- Abstract(参考訳): VLA(Vision-Language-Action)ファウンデーションモデルは、視覚言語プレトレーニング中に得られた知識を正確かつ解釈可能な運転に活用するため、近年、自律運転の一般的なソリューションの1つとして現れている。
しかし、VLAは画像のトークンコストが高く、全方向停止時の到着順序の決定や長距離運転シーン理解などのメモリ依存タスクでは問題となるため、視覚的な入力として限られたフレームしか持たない。
既存のソリューションでは、クロスアテンションを通じてアクセスされる潜在ベクトルメモリを使用し、解釈可能でもポータブルでもない。
本稿では,言語ベースのメモリを備えた汎用VLA駆動エージェントAD-Memoを提案する。
エージェントは、そのChain-of-Thought(CoT)の拡張としてメモリを出力し、運転に不可欠な周囲のオブジェクトを記録する。
Supervised Fine-Tuning (SFT) と \textit{Da Capo} という新しい半閉鎖ループ強化学習 (RL) アルゴリズムは、メモリに対するトラジェクトリレベルの利点と、ドライブに対するステップレベルのアドバンテージを利用しており、より優れたクレジット割り当てをもたらす。
オールウェイ停止や一般運転といったシナリオを越えて、AD-Memoは運転品質を改善し、運転シーンでの質問応答を改善するとともに、他のモデルのプラグアンドプレイメモリを生成する。
関連論文リスト
- Rethinking Language's Role in Efficient VLA for Autonomous Vehicles: Toward Smarter, Trustworthy Driving [8.920589816043298]
Vision-Language-Action(VLA)モデルが自動走行(AD)を改造
レイテンシとメモリの予算は厳しく、自動回帰デコーディングは本質的にシーケンシャルです。
この作業は、トレーニングコストが1回だけ支払われる間に、デプロイされたフレーム毎に推論コストが再帰するため、言語がいつ、いつ、どこで振る舞うべきかという中心的な疑問を再考する。
論文 参考訳(メタデータ) (2026-08-31T01:51:51Z) - Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation [96.18832998704329]
ラメムVLA(LaMem-VLA)は、ラメムVLA(LaMem-VLA)という、ラメムVLA(LaMem-VLA)という、ラメムVLA(LaMem-VLA)という、ラメムVLA(LaMem-VLA)という、ラメムVLA(LaMem-VLA)のフレームワークである。
LaMem-VLAは、メモリがVLA推論に直接参加し、コンテキスト境界下でアクション生成をガイドすることを可能にする。
論文 参考訳(メタデータ) (2026-07-08T16:26:06Z) - AutoMem: Automated Learning of Memory as a Cognitive Skill [63.07128029793901]
記憶の専門知識は学習したスキルであることを示す。
メモリ管理の2つの軸を自動化するフレームワークであるAutoMemを紹介する。
その結果,メモリ管理は独立して学習可能なスキルであることが示唆された。
論文 参考訳(メタデータ) (2026-07-01T17:57:03Z) - Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning [89.55738101744657]
大規模言語モデル(LLM)は、幅広いNLPタスクで印象的な機能を示しているが、基本的にはステートレスである。
本稿では,LLMに外部メモリを積極的に管理・活用する機能を備えた強化学習フレームワークであるMemory-R1を提案する。
論文 参考訳(メタデータ) (2025-08-27T12:26:55Z) - FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning [75.80110543049783]
我々は,自律運転のための再建型視覚トークンプルーニングフレームワークであるFastDriveVLAを提案する。
VLAモデルの視覚的エンコーダにReconPrunerを訓練するために, 新たなフォアグラウンド逆バックグラウンド再構築戦略を考案した。
提案手法は,異なるプルーニング比におけるnuScenesオープンループ計画ベンチマークの最先端結果を実現する。
論文 参考訳(メタデータ) (2025-07-31T07:55:56Z) - Multi-Frame, Lightweight & Efficient Vision-Language Models for Question Answering in Autonomous Driving [0.0]
我々は,自律運転のための視覚質問応答を行う,効率的で軽量な多フレーム視覚言語モデルを開発した。
従来のアプローチと比較して、EM-VLM4ADは少なくとも10倍のメモリと浮動小数点演算を必要とする。
論文 参考訳(メタデータ) (2024-03-28T21:18:33Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。