論文の概要: ActiveArena: Benchmarking and Understanding Active Perception in Robotic Manipulation
- arxiv url: http://arxiv.org/abs/2609.24124v2
- Date: Wed, 23 Sep 2026 05:53:55 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-25 00:05:17.668612
- Title: ActiveArena: Benchmarking and Understanding Active Perception in Robotic Manipulation
- Title(参考訳): ActiveArena: ロボットマニピュレーションにおけるアクティブ知覚のベンチマークと理解
- Abstract要約: ActiveArena-Simは、制御可能な視点と大規模なワークスペースを備えたアクティブパーセプションシミュレータである。
ActiveArena-Benchは5つの細かいカテゴリにわたる35のタスクで構成され、視覚的な探索とインタラクティブな情報取得をカバーしている。
ActiveArena-VLAは、メモリ書き込み、メモリ容量、受容状態、サブタスクの監督、アクティブな知覚における高レベル計画の制御のための13の視覚言語対応構成のモジュールスイートである。
- 参考スコア(独自算出の注目度): 46.7774490610328
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Active perception and manipulation are crucial for robots to interact with complex scenes. Existing benchmarks struggle to evaluate how robots effectively acquire and maintain information in memory in an active manner. To this end, we introduce ActiveArena-Sim, an active-perception simulator with controllable viewpoints and large-scale workspaces as the foundation. Built on this, we propose ActiveArena-Bench, which comprises 35 tasks across 5 fine-grained categories, covering visual exploration and interactive information acquisition. Each task is difficult to solve from passive observations alone, requiring multi-round evidence acquisition and memory-based reasoning. The benchmark provides rich memory annotations, standardized training data, and ID/OOD protocols featuring disjoint scenes, unseen distractor configurations, and novel backgrounds. Moreover, we present ActiveArena-VLA, a modular suite of 13 vision-language-action configurations for controlled studies of memory writing, memory capacity, proprioceptive state, subtask supervision, and high-level planning in active perception. Benchmark results reveal a substantial ID-OOD gap: uniform memory sampling, increased memory capacity under reliable write policies, proprioceptive inputs, and subtask supervision improve OOD generalization, while planner-guided memory management and decision-making achieve performance close to the best-performing configuration using only sparse memory. ActiveArena thus provides a unified testbed to develop and diagnose models for active perception and manipulation.
- Abstract(参考訳): ロボットが複雑なシーンと対話するためには、アクティブな知覚と操作が不可欠である。
既存のベンチマークでは、ロボットがメモリ内の情報をアクティブに取得し、維持する方法を評価するのに苦労している。
この目的のために,制御可能な視点と大規模ワークスペースを備えたアクティブ知覚シミュレータであるActiveArena-Simを紹介した。
そこで本研究では,視覚探索と対話型情報取得を対象とする,5つの細かなカテゴリにわたる35のタスクからなるActiveArena-Benchを提案する。
各タスクは受動的観察だけでは解決が困難であり、複数ラウンドのエビデンス取得とメモリベースの推論が必要となる。
このベンチマークは、リッチなメモリアノテーション、標準化されたトレーニングデータ、および相容れないシーン、目に見えないイントラクタ構成、新しいバックグラウンドを含むID/OODプロトコルを提供する。
さらに、メモリ書き込み、メモリキャパシティ、受容状態、サブタスクの監視、アクティブな知覚における高レベル計画の制御のための13の視覚言語アクション構成のモジュールスイートであるActiveArena-VLAを提案する。
メモリ一様サンプリング、信頼性の高い書き込みポリシー下でのメモリ容量の増加、プロプリセプティブインプット、サブタスクインスペクションによりOODの一般化が向上し、プランナー誘導型メモリ管理と意思決定はスパースメモリのみを使用して最高のパフォーマンス設定に近いパフォーマンスを達成する。
これによりActiveArenaは、アクティブな知覚と操作のためのモデルの開発と診断のための統合テストベッドを提供する。
関連論文リスト
- MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks [55.145729491377374]
メモリを持つエージェントの既存の評価は、通常、単独で記憶と行動を評価する。
マルチセッションメモリ-エージェント環境ループにおけるエージェントメモリのベンチマークのための統合評価ジムであるMemoryArenaを紹介する。
MemoryArenaは、Webナビゲーション、優先制約付き計画、プログレッシブ情報検索、シーケンシャルなフォーマルな推論を含む評価をサポートする。
論文 参考訳(メタデータ) (2026-02-18T09:49:14Z) - ActMem: Bridging the Gap Between Memory Retrieval and Reasoning in LLM Agents [14.695250837875454]
本稿では,ActMemと呼ばれる新しい動作可能なメモリフレームワークを提案する。
ActMemは非構造化対話履歴を構造化因果グラフと意味グラフに変換する。
エージェントは暗黙の制約を推論し、過去の状態と現在の意図の間の潜在的な衝突を解決することができる。
論文 参考訳(メタデータ) (2026-02-04T00:54:53Z) - MineNPC-Task: Task Suite for Memory-Aware Minecraft Agents [7.941984883391391]
我々は,MineNPC-Taskを紹介した。これは,オープンワールドMinecraftにおけるメモリ認識,混合開始LDMエージェントのテストのための,ユーザ認可ベンチマークと評価ハーネスである。
合成プロンプトに頼るのではなく、タスクはプロのプレイヤーと形式的で要約的なコプレイによって引き起こされ、その後パラメトリックテンプレートに正規化される。
ハーネスはプラン、アクション、メモリイベントをキャプチャし、プランプレビュー、対象の明確化、メモリ読み込みと書き込み、プレコンディションチェック、修復の試みを含む。
論文 参考訳(メタデータ) (2026-01-08T18:39:52Z) - FindingDory: A Benchmark to Evaluate Memory in Embodied Agents [49.18498389833308]
本研究では,Habitatシミュレータに長距離エンボディタスクのための新しいベンチマークを導入する。
このベンチマークは、持続的なエンゲージメントとコンテキスト認識を必要とする60タスクにわたるメモリベースの機能を評価する。
論文 参考訳(メタデータ) (2025-06-18T17:06:28Z) - Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO [63.140883026848286]
アクティブビジョン(Active Vision)とは、タスク関連情報を収集するために、どこでどのように見るべきかを積極的に選択するプロセスである。
近年,マルチモーダル大規模言語モデル (MLLM) をロボットシステムの中心的計画・意思決定モジュールとして採用する動きが注目されている。
論文 参考訳(メタデータ) (2025-05-27T17:29:31Z) - How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior [65.70584076918679]
メモリは、大きな言語モデル(LLM)ベースのエージェントにおいて重要なコンポーネントである。
本稿では,メモリ管理の選択がLLMエージェントの行動,特に長期的パフォーマンスに与える影響について検討する。
論文 参考訳(メタデータ) (2025-05-21T22:35:01Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。