論文の概要: Unified Condition-Action Modeling for Accurate One-Step Action Generation
- arxiv url: http://arxiv.org/abs/2608.16153v3
- Date: Sun, 23 Aug 2026 16:38:01 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-25 18:24:36.877152
- Title: Unified Condition-Action Modeling for Accurate One-Step Action Generation
- Title(参考訳): 高精度ワンステップ動作生成のための統一条件-アクションモデリング
- Authors: Xinyu Zhou, Zikun Cai, Kuangji Zuo, Gen Li, Boyu Ma, Yanshuo Lu, Yutong Song, Mingqi Yuan, Jiayu Chen, Jianfei Yang,
- Abstract要約: UCA-Flowは、正確なワンステップアクション生成のための統一された条件-アクションモデリングフレームワークである。
本手法は,観測条件,時間経過条件,間隔条件,アクショントークンを単一シーケンスに統一し,統一条件-アクション変換器で処理する。
- 参考スコア(独自算出の注目度): 24.554218523494498
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Robot manipulation requires policies that are both accurate and efficient, as robot control must respond to changing observations under tight latency constraints. Recent diffusion and flow policies are promising, but they often treat conditions as auxiliary signals rather than jointly evolving them with action trajectories. We find that this limitation can be effectively mitigated by a \textbf{simple yet effective unified condition-action modeling design} that represents conditions and actions in a shared token space, allowing a compact model to achieve high performance while improving both inference speed and accuracy. Therefore, we propose UCA-Flow, a unified condition-action modeling framework for accurate one-step action generation. Our method unifies observation conditions, timestep conditions, interval conditions, and action tokens into a single sequence, and processes them with a Unified Condition-Action Transformer for joint condition-action representation learning. As a result, condition representations are dynamically reconstructed according to the current generation stage, highlighting information most relevant for action refinement. Furthermore, we introduce an improved dual-pass supervision scheme over $u$ and $v$ for stronger optimization of unified condition-action modeling. UCA-Flow improves the average success rate by 9.3 percentage points over the strongest baseline, while achieving $45.6\times$ and $33.4\times$ speedups over DP3 and Simple DP3, and remaining $4.3\times$ and $2.3\times$ faster than one-step FlowPolicy and MP1, respectively.
- Abstract(参考訳): ロボットの操作には正確かつ効率的なポリシーが必要である。
近年の拡散と流路政策は有望であるが、しばしば行動軌跡を共同で進化させるのではなく、補助的な信号として扱う。
この制限は、共有トークン空間における条件や動作を表現するための「textbf{simple yet effective unified condition-action modeling design」によって効果的に緩和され、コンパクトモデルが推論速度と精度の両方を改善しながら高い性能を達成することができる。
そこで本研究では,正確なワンステップアクション生成のための統一された条件-アクションモデリングフレームワークであるUCA-Flowを提案する。
本手法は,観測条件,時間経過条件,間隔条件,行動トークンを単一シーケンスに統一し,共同条件-動作表現学習のための統一条件-動作変換器で処理する。
その結果、状態表現は、現在の生成段階に応じて動的に再構成され、アクションリファインメントに最も関連性の高い情報をハイライトする。
さらに、統一された条件-動作モデリングのより強力な最適化のために、$u$と$v$で改良されたデュアルパス監視方式を導入する。
UCA-Flowは最強のベースラインで平均成功率を9.3%向上させ、45.6\times$と33.4\times$はDP3とSimple DP3を上回り、残り4.3\times$と2.3\times$はそれぞれ1ステップのFlowPolicyとMP1より速い。
関連論文リスト
- PIER-Flow: Physics-Informed Efficient Rectified Flow for Real-Time Mobile Robot Navigation [10.164953732323463]
PIER-Flowは移動ロボットのための軽量ナビゲーションポリシーである。
98.85%の成功率とゼロ衝突を達成し、平均値は$sim$1.29 msである。
論文 参考訳(メタデータ) (2026-07-11T12:34:28Z) - SA-VLA: State-aware tokenizer for improving Vision-Language-Action Models' performance [12.586771228417112]
本稿では,動作復号をロボットの状態に設定した状態認識型アクショントークンであるSA-VLAを提案する。
12のRoboTwin操作タスクでは、SA-VLAは最強のトークン化剤ベースラインよりも平均成功率を0.29から0.56に改善する。
論文 参考訳(メタデータ) (2026-06-29T10:45:53Z) - PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models [50.232333672172395]
VLA(Vision-Language-Action)モデルは、ロボット操作の統一パラダイムを提供するが、実際のデプロイメントは実行効率によってボトルネックになることが多い。
強化学習に基づくポストトレーニングフレームワークである textbfPolicyTrim を提案する。
私たちのフレームワークは、タスクの成功率を損なうことなく、最大5.83$times$エンドツーエンドのデプロイメントスピードアップを提供します。
論文 参考訳(メタデータ) (2026-06-21T14:54:07Z) - Dynamic Execution Commitment of Vision-Language-Action Models [21.647844049489535]
本稿では,動的実行コミットメントを自己特定的プレフィックス検証問題として再編成する適応アクションアクセプタンス機構であるA3を紹介する。
A3はまず、グループサンプリングを介して行動の軌跡的なコンセンサススコアを計算し、次に代表ドラフトを選択し、下流検証を優先する。
さまざまなVLAモデルとベンチマークの実験では、A3は手動の水平調整の必要性を排除し、実行と推論のスループットのトレードオフを優れたものにしている。
論文 参考訳(メタデータ) (2026-05-12T05:52:58Z) - Efficient Hybrid SE(3)-Equivariant Visuomotor Flow Policy via Spherical Harmonics for Robot Manipulation [51.48086451600577]
E3Flowは、同変拡散ポリシーの限界に対処する新しいフレームワークである。
安定な多モード同変学習による効率的な整流を初めて統一する。
E3Flowは、最先端の球拡散政策よりも平均的な成功率を3.12%向上させる。
論文 参考訳(メタデータ) (2026-03-24T14:00:36Z) - Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action Generation [65.13627721310613]
平均速度ポリシー(MVP)は、平均速度場をモデル化し、最速のワンステップアクション生成を実現するための新しい生成ポリシー関数である。
MVPはRoomimicとOGBenchのいくつかの困難なロボット操作タスクに対して、最先端の成功率を達成する。
論文 参考訳(メタデータ) (2026-02-14T14:44:06Z) - ReSeFlow: Rectifying SE(3)-Equivariant Policy Learning Flows [7.360373380580255]
本稿では, SE(3)-拡散モデルに補正を導入し, 高速かつジオデシックな, 最小計算型ポリシー生成を提供するReSeFlowを提案する。
提案したReSeFlowは,提案手法よりも測地距離が低い場合に高い性能が得られることがわかった。
論文 参考訳(メタデータ) (2025-09-20T06:32:36Z) - SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model Acceleration [70.72227437717467]
VLA(Vision-Language-Action)モデルは、その強力な制御能力に注目が集まっている。
計算コストが高く、実行頻度も低いため、ロボット操作や自律ナビゲーションといったリアルタイムタスクには適さない。
本稿では,共同スケジューリングモデルとプルーニングトークンにより,VLAモデルを高速化する統一フレームワークSP-VLAを提案する。
論文 参考訳(メタデータ) (2025-06-15T05:04:17Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。