論文の概要: Task-Adaptive Grounded 3D-Programmers Using 2D VLMs
- arxiv url: http://arxiv.org/abs/2610.02021v1
- Date: Thu, 01 Oct 2026 16:43:17 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-03 01:19:24.304616
- Title: Task-Adaptive Grounded 3D-Programmers Using 2D VLMs
- Title(参考訳): 2次元VLMを用いたタスク適応型接地型3次元プログラマー
- Abstract要約: CCF(Canonical Coordinate Framing)とTAF(Task-Adaptive Feedback)の2つの新しい概念で、3Dグラウンドと反復的なフィードバックループを導入する。
CCFは、共通のユークリッド座標系に入力と出力の両方をアンカーする統一された視覚表現として機能する。
TAFは、タスク適応動的フィードバックで推論ループをクローズし、2次元のVLMがネイティブな視覚的コンテキスト内で様々なオープン語彙タスクを実行できるようにする。
- 参考スコア(独自算出の注目度): 49.104379630333916
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by data scale, training diversity, and reasoning capacity. Instead of naively extending these models into 3D, we take a different approach: we enable powerful 2D VLMs to operate reliably in 3D by introducing 3D grounding and iterative feedback loops with two novel concepts: Canonical Coordinate Framing (CCF) and Task-Adaptive Feedback (TAF). CCF serves as a unified visual representation that anchors both inputs and outputs to a shared Euclidean coordinate system, solving common challenges in 3D grounding such as axis ambiguity, inconsistent metric scale, and floating references. Complementary to this structured framing of the 3D inputs, TAF closes the reasoning loop with task-adaptive dynamic feedback that enables 2D VLMs to perform varied open-vocabulary tasks within their native visual context. Building on this foundation, we introduce 3D-Prog, a 3D understanding, reasoning, and generation framework that jointly employs the capabilities of CCF and TAF together with powerful VLMs. Without requiring any retraining, 3D-Prog performs open-vocabulary 3D understanding, manipulation, and generation across both object-level and scene-level tasks. Our experiments show that the joint use of CCF and TAF transforms 2D VLMs into geometry-aware 3D programmers, achieving consistent, interpretable, and high-quality results across diverse 3D tasks.
- Abstract(参考訳): 近年の視覚言語モデル (VLM) は、顕著な一般化と推論能力を示すが、これらのモデルにおける3次元理解は、データスケール、訓練の多様性、推論能力によって制限されている。
強力な2D VLMを3Dで確実に動作させるには、CCF(Canonical Coordinate Framing)とTAF(Task-Adaptive Feedback)という2つの新しい概念の3Dグラウンドと反復的なフィードバックループを導入する必要がある。
CCFは、入力と出力の両方を共用ユークリッド座標系に固定する統一された視覚表現として機能し、軸のあいまいさ、不整合計量スケール、浮動小数点参照などの3Dグラウンドにおける共通の課題を解決する。
3D入力の構造的フレーミングと相まって、TAFは2D VLMがネイティブな視覚的コンテキスト内で様々なオープン語彙タスクを実行できるタスク適応動的フィードバックで推論ループを閉じる。
この基盤の上に構築された3D-Progは,CCFとTAFの能力と強力なVLMを併用した3D理解・推論・生成フレームワークである。
再トレーニングを必要としない3D-Progは、オブジェクトレベルとシーンレベルの両方のタスクに対してオープンな3D理解、操作、生成を行う。
実験の結果,CCFとTAFの併用により2次元VLMを幾何学的に認識された3次元プログラマに変換し,多種多様な3次元タスクに対して一貫した,解釈可能な,高品質な結果が得られることがわかった。
関連論文リスト
- ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA [26.947027118604368]
textbfViewMind3Dは、シーンを多視点で観察する3次元空間推論のための、完全にトレーニング不要でモジュラーなフレームワークである。
ViewMind3Dは,事前のトレーニング不要および微調整3D-LLMと比較して,競争性能が向上することを示す。
論文 参考訳(メタデータ) (2026-07-30T16:14:30Z) - Lift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware Manipulation [76.16065370700488]
We introduced Lift3D-VLA, a unified VLA framework that equips models with explicit 3D point cloud reasoning and possible temporally coherent action generation。
我々は,Lift3D-VLAがMetaWorldとRLBenchの平均成功率を10.8%,11.1%向上したことを示す。
論文 参考訳(メタデータ) (2026-07-07T17:59:47Z) - 3DThinkVLA: Endowing Vision-Language-Action Models with Latent 3D Priors via 3D-Thinking-Guided Co-training [19.5808550016589]
3次元幾何学的知覚と3次元空間的推論は、異なる特徴階層で切り離され、注入される異なる能力である。
本稿では,視覚言語行動モデル(VLA)が行動予測中に暗黙的に3次元空間推論を行うことを可能にする3次元思考誘導協調学習フレームワークを提案する。
論文 参考訳(メタデータ) (2026-06-03T04:34:07Z) - Abstract 3D Perception for Spatial Intelligence in Vision-Language Models [100.13033631690114]
視覚言語モデル(VLM)は、空間認識や物理的理解といった3D関連課題に苦しむ。
我々は,VLMの幾何学的構造と物理力学を符号化するために,抽象的境界ボックスを利用するフレームワークであるSandboxVLMを紹介した。
提案手法は空間知能を常に向上させ,SAT Realの8.3%のゲインをベースライン法と比較して達成する。
論文 参考訳(メタデータ) (2025-11-14T04:16:09Z) - 3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation [17.294440057314812]
VLM(Vision-Language Models)は様々な視覚的・言語的タスクにおいて顕著な性能を示した。
人為的な幾何学的手がかりを予め訓練されたVLMに注入するフレームワークであるGeometric Distillationを提案する。
本手法は、自然な画像テキスト入力と互換性を保ちながら、表現を幾何学的に認識するように形成する。
論文 参考訳(メタデータ) (2025-06-11T15:56:59Z) - VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction [86.82819259860186]
本稿では,視覚言語モデル(VLM)のための統合フレームワークであるVLM-3Rについて紹介する。
VLM-3Rは、空間的理解を表す暗黙の3Dトークンを導出する幾何学エンコーダを用いて、モノクロビデオフレームを処理する。
論文 参考訳(メタデータ) (2025-05-26T17:56:30Z) - Multi-CLIP: Contrastive Vision-Language Pre-training for Question
Answering tasks in 3D Scenes [68.61199623705096]
一般的な言語知識と視覚概念を2次元画像から3次元シーン理解に適用するためのトレーニングモデルは、研究者が最近探求を始めたばかりの有望な方向である。
そこで本研究では,モデルによる3次元シーンポイントクラウド表現の学習を可能にする,新しい3次元事前学習手法であるMulti-CLIPを提案する。
論文 参考訳(メタデータ) (2023-06-04T11:08:53Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。