論文の概要: AnyStep-WAM: Budget-Aligned Distillation and Adaptive Inference for World Action Models
- arxiv url: http://arxiv.org/abs/2609.33748v1
- Date: Sun, 27 Sep 2026 16:49:05 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-01 18:57:25.890992
- Title: AnyStep-WAM: Budget-Aligned Distillation and Adaptive Inference for World Action Models
- Title(参考訳): AnyStep-WAM:世界行動モデルに対する予算適応蒸留と適応推論
- Abstract要約: 我々は、調整可能な予算予測とシーン依存アロケーションのための一般的なフレームワークであるAnyStep World Action Modelを紹介する。
軽量なリスクベネフィットスケジューラは、1段階のプレビューから難易度と予算固有の生徒教育者の忠実度を予測する。
本手法は, タスク成功率を維持しつつ, 平均復調歩数を60.2%, 49.8%, 85.28%削減する。
- 参考スコア(独自算出の注目度): 31.57475195690971
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: World-action models (WAMs) couple predictive visual modeling with action generation, typically relying on iterative denoising with a fixed denoising steps. However, manipulation tasks contain actions chunks with varying sensitivity to generation errors: critical actions require precision, while less sensitive actions allow faster generation with fewer denoising steps. Here we introduce AnyStep World Action Model, a general framework for tunable-budget prediction and scene-dependent computation allocation. Our budget-aligned teacher-trajectory distillation trains interval-conditioned flow maps using explicit frozen-teacher transitions and shared low-rank adapters, supporting action generation from one-step prediction to multi-step refinement. Building on this capability, a lightweight risk-benefit scheduler predicts teacher-curvature-based difficulty and budget-specific student-teacher fidelity from a single one-step preview, selecting the smallest budget predicted to satisfy risk-adaptive fidelity requirements. We evaluate our framework on three widely used WAMs Motus, FastWAM, and LingBotVA using RoboTwin 2.0. Our method reduces average denoising steps by 60.2%, 49.8%, and 85.28%, respectively, while maintaining baseline task success rates. In particular, our AnyStep training substantially improves model performance under a one-step denoising budget, increasing task success rates by 7.07%, 12.08%, and 8.94% on Motus, FastWAM, and LingBotVA, respectively. Experiments on six real-world manipulation tasks further validate its effectiveness.
- Abstract(参考訳): ワールドアクションモデル(WAM)は、予測的視覚モデルとアクション生成を結合する。
しかし、操作タスクには、エラー発生に対する感度の異なるアクションチャンクが含まれている。
ここでは、調整可能な予算予測とシーン依存の計算割り当てのための一般的なフレームワークであるAnyStep World Action Modelを紹介する。
具体的凍結型ティーチング・トランジションと低ランクアダプタの共用による教員軌道蒸留列車のインターバルコンディショニング・フローマップを試作し,一段階予測から多段階改良までのアクション生成を支援した。
この能力に基づいて、軽量なリスクベネフィットスケジューラは、1段階のプレビューから教師の曲率に基づく難易度と予算固有の生徒-教師の忠実度を予測し、リスク適応フィデリティ要件を満たすために予測される最小の予算を選択する。
我々は,RoboTwin 2.0を用いて,Motus,FastWAM,LingBotVAの3つのWAMについて検討を行った。
本手法は, タスク成功率を維持しつつ, 平均復調歩数を60.2%, 49.8%, 85.28%削減する。
特に、AnyStepトレーニングは、1段階のデノベーション予算の下でモデルパフォーマンスを大幅に向上させ、タスク成功率を7.07%、12.08%、Motus、FastWAM、LingBotVAで8.94%向上させました。
6つの実世界の操作タスクの実験は、その効果をさらに検証した。
関連論文リスト
- Denoising Tells When to Replan: Denoising-Variance Adaptive Chunking for Flow-Based Robot Policies [18.476299874941592]
アクションチャンキングはフローベースのロボットポリシーの一般的な推論戦略となっている。
本研究では,予測チャンクから実行すべきアクション数を適応的に決定するテストタイム手法であるDVACを提案する。
LIBERO、RoboTwin、CALVIN、および実世界の操作実験により、DVACはスケジュール変更頻度を減らしながらタスク成功を改善することが示された。
論文 参考訳(メタデータ) (2026-06-02T16:26:32Z) - A1: A Fully Transparent Open-Source, Adaptive and Efficient Truncated Vision-Language-Action Model [112.9420001646428]
VLA(Vision-Language-Action)モデルは、オープンワールドロボット操作の強力なパラダイムとして登場したが、実際の展開はコストに制約されることが多い。
我々は、低コストで高スループットな推論のために設計された、完全にオープンソースで透明なVLAフレームワークであるA1を提示する。
A1は最先端の成功率を達成すると同時に、推論コストを大幅に削減する。
論文 参考訳(メタデータ) (2026-04-07T10:18:40Z) - Sculpting Subspaces: Constrained Full Fine-Tuning in LLMs for Continual Learning [19.27175827358111]
大規模言語モデル(LLM)における継続的な学習は破滅的な忘れがちである。
適応特異値分解(SVD)を利用した連続的完全微調整手法を提案する。
我々は,Encoder-decoder (T5-Large) モデルとdecoder-only (LLaMA-2 7B) モデルの両方を用いて,標準連続学習ベンチマークを広範囲に評価した。
論文 参考訳(メタデータ) (2025-04-09T17:59:42Z) - Adaptive Training Meets Progressive Scaling: Elevating Efficiency in Diffusion Models [52.1809084559048]
TDCトレーニングと呼ばれる新しい2段階分割型トレーニング戦略を提案する。
タスクの類似性と難易度に基づいてタイムステップをグループ化し、高度にカスタマイズされた復調モデルを各グループに割り当て、拡散モデルの性能を向上させる。
2段階のトレーニングでは、各モデルを個別にトレーニングする必要がなくなるが、総トレーニングコストは、単一の統合されたデノナイジングモデルをトレーニングするよりもさらに低い。
論文 参考訳(メタデータ) (2023-12-20T03:32:58Z) - Towards Practical Lipreading with Distilled and Efficient Models [57.41253104365274]
ニューラルネットワークの復活により、リリーディングは多くの進歩を目の当たりにした。
最近の研究は、最適なアーキテクチャを見つけるか、一般化を改善することで、パフォーマンスを改善するといった側面に重点を置いている。
現在の方法論と、実践的なシナリオにおける効果的なリップリーディングのデプロイ要件との間には、依然として大きなギャップがあります。
まず, LRW と LRW-1000 をそれぞれ 88.5% と 46.6% に比例して, 最先端の性能を高めることを提案する。
論文 参考訳(メタデータ) (2020-07-13T16:56:27Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。