論文の概要: GPT-6-Astra Lights Up Embodied Navigation: Evaluation in Zero-Shot Vision-and-Language Navigation in Continuous Environments
- arxiv url: http://arxiv.org/abs/2609.29861v2
- Date: Fri, 25 Sep 2026 12:44:47 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-28 18:27:51.057901
- Title: GPT-6-Astra Lights Up Embodied Navigation: Evaluation in Zero-Shot Vision-and-Language Navigation in Continuous Environments
- Title(参考訳): GPT-6-Astra Lights Up Embodied Navigation: In Zero-Shot Vision-and-Language Navigation in Continuous Environments
- Abstract要約: 我々は,汎用基盤モデルであるGPT-6-Astraが,自覚,推論,意思決定能力を用いて,慣れていない環境をナビゲートできるかどうかを検討する。
単分子RGBを用いて、GPT-6-Astraはいつ観測し、どのように移動し、いつ停止するかを決定する。
R2R-CE-100では、GPT-6-Astraが81.3%の成功率を記録し、最強のゼロショットと、列車による成功率も15.3ポイント、9.2ポイントを超えた。
- 参考スコア(独自算出の注目度): 43.03393333823919
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: We investigate whether GPT-6-Astra, a general-purpose foundation model, can navigate unfamiliar environments using its own perception, reasoning, and decision-making capabilities. Our evaluation focuses on zero-shot vision-and-language navigation in continuous environments (VLN-CE) through a minimal interface in the Codex harness, aiming to unleash GPT-6-Astra's full potential for navigation. Using monocular RGB, GPT-6-Astra decides when to observe, how to move, and when to stop, without navigation-specific fine-tuning, a trained waypoint predictor, or a pre-built scene map. Our evaluation yields four key findings and implications. First, GPT-6-Astra achieves strong zero-shot navigation performance using only monocular RGB observations. On R2R-CE-100, GPT-6-Astra (ultra reasoning) achieves a success rate of 81.3%, exceeding the strongest reported zero-shot and even train-based success rates by 15.3 and 9.2 percentage points, respectively. Second, GPT-6-Astra exhibits promising capabilities in interpreting multi-stage instructions, understanding the environment, and adjusting routes. Third, execution and goal-verification failures persist even with ultra reasoning. Fourth, these results motivate combining general-purpose model capabilities with navigation-specific expertise. Based on these findings, future VLN research should investigate which aspects of instruction interpretation, spatial understanding, and navigation decision-making general-purpose models can handle directly, and where navigation-specific learning can extend their capabilities. This includes exploring how spatial representations, navigation experience, and learned control skills can improve progress tracking, error recovery, and goal verification while preserving the flexibility to adjust routes.
- Abstract(参考訳): 我々は,汎用基盤モデルであるGPT-6-Astraが,自覚,推論,意思決定能力を用いて,慣れていない環境をナビゲートできるかどうかを検討する。
我々は,連続環境(VLN-CE)におけるゼロショット・ヴィジュアル・アンド・ランゲージナビゲーションを,GPT-6-Astraのナビゲーションの可能性を完全に解き放つことを目的として,Codexハーネス内の最小限のインタフェースを用いて評価した。
GPT-6-Astraは、単眼のRGBを使用して、ナビゲーション固有の微調整、訓練されたウェイポイント予測器、または事前に構築されたシーンマップなしで、いつ、どのように動くか、いつ停止するかを決定する。
我々の評価は4つの重要な発見と意味をもたらす。
まず、GPT-6-Astraは単眼RGB観測のみを用いて強力なゼロショットナビゲーション性能を実現する。
R2R-CE-100では、GPT-6-Astra (ultra reasoning) が81.3%の成功率に達し、最強のゼロショット、さらには列車ベースの成功率も15.3および9.2ポイントである。
第2に、GPT-6-Astraは多段階命令の解釈、環境の理解、経路の調整において有望な能力を示す。
第三に、実行とゴール検証の失敗は、極端に推論しても継続する。
第4に、これらの結果は、汎用モデル機能とナビゲーション特有の専門知識の組み合わせを動機付けている。
これらの知見に基づき、今後のVLN研究は、命令解釈、空間的理解、ナビゲーション決定汎用モデルのどの側面が直接処理できるのか、ナビゲーション固有の学習が機能を拡張することができるのかを調査すべきである。
これには、空間表現、ナビゲーション経験、学習した制御スキルが、ルートを調整する柔軟性を維持しながら、進捗追跡、エラー回復、目標検証を改善する方法についての調査が含まれる。
関連論文リスト
- How Far Can GPT-6-Astra Go? Evaluating Capabilities in Zero-Shot Vision-and-Language Navigation [20.65686859199639]
連続環境におけるゼロショットビジョン・アンド・ランゲージナビゲーションにおけるGPT-6-Astraについて検討する。
指示を解釈し、その環境を評価し、行動を提案する。
Open-Nav が使用した100R2R-CE val-unseen エピソードの50件について,本システムの評価を行った。
論文 参考訳(メタデータ) (2026-09-17T12:14:43Z) - LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation [52.91080323083409]
エージェントは異質な目標や視覚的な観察をタスク、環境、ロボットの体現物に翻訳する必要がある。
ここでは,タスク固有の予測ヘッドを使わずに,予め訓練されたVLMの空間的知性を取り入れた,コンパクトな一般化型ナビゲーションモデルLightNav-0を提案する。
論文 参考訳(メタデータ) (2026-08-31T15:08:53Z) - Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation [5.156481604986546]
本稿では,Embodied-R1上に構築された隠れ状態空間読み出し型Pointing-VLAを提案する。
Pointing-VLA は Bridge/WidowX 上で SOTA のパフォーマンスを達成する。
論文 参考訳(メタデータ) (2026-08-24T11:43:49Z) - ABot-N1: Toward a General Visual Language Navigation Foundation Model [27.085429358326504]
ABot-N1は、一般的なVisual Language Navigation基盤モデルへのステップである。
スロービジョン言語推論器は、画素ゴールを生成しながら明示的なChain-of-Thought推論を行う。
新しいベンチマークは、都市規模のナビゲーションの分野を前進させるためにオープンソースとしてリリースされている。
論文 参考訳(メタデータ) (2026-07-11T16:21:03Z) - Correctable Landmark Discovery via Large Models for Vision-Language Navigation [89.15243018016211]
Vision-Language Navigation (VLN) は、ターゲット位置に到達するために、エージェントが言語命令に従う必要がある。
以前のVLNエージェントは、特に探索されていないシーンで正確なモダリティアライメントを行うことができない。
我々は,Large ModEls (CONSOLE) によるコレクタブルLaNdmark DiScOveryと呼ばれる新しいVLNパラダイムを提案する。
論文 参考訳(メタデータ) (2024-05-29T03:05:59Z) - NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning [97.88246428240872]
Embodied AIの重要な研究課題であるVision-and-Language Navigation (VLN)は、自然言語の指示に従って複雑な3D環境をナビゲートするために、エンボディエージェントを必要とする。
近年の研究では、ナビゲーションの推論精度と解釈可能性を改善することにより、VLNにおける大きな言語モデル(LLM)の有望な能力を強調している。
本稿では,自己誘導型ナビゲーション決定を実現するために,パラメータ効率の高いドメイン内トレーニングを実現する,Navigational Chain-of-Thought (NavCoT) という新しい戦略を提案する。
論文 参考訳(メタデータ) (2024-03-12T07:27:02Z) - Rethinking the Spatial Route Prior in Vision-and-Language Navigation [29.244758196643307]
VLN(Vision-and-Language Navigation)は、知的エージェントを自然言語による予測位置へナビゲートすることを目的としたトレンドトピックである。
この研究は、VLNのタスクを、これまで無視されていた側面、すなわちナビゲーションシーンの前の空間ルートから解決する。
論文 参考訳(メタデータ) (2021-10-12T03:55:43Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。