論文の概要: Guided Discovery of New Behaviors using Diffusion Policies
- arxiv url: http://arxiv.org/abs/2606.08743v1
- Date: Sun, 07 Jun 2026 17:24:44 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-09 14:42:06.426442
- Title: Guided Discovery of New Behaviors using Diffusion Policies
- Title(参考訳): 拡散反応を利用した新しい行動の探索
- Authors: Dian Yu, Sebastian Sanokowski, Majid Khadiv,
- Abstract要約: 本稿では,拡散政策のサンプルを将来性に乏しいサンプルへ誘導する枠組みを提案する。
提案手法は, 新規な軌道を効果的にマイニングし, 修復し, 多様な, 実行可能な動作の体系的な発見を可能にする。
- 参考スコア(独自算出の注目度): 10.086845726616694
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Diffusion models have become a powerful tool for generative modeling in robotics, with diffusion policies excelling at modeling multimodal action-trajectory distributions. However, when demonstrations are limited, standard sampling often reproduces dominant behaviors while neglecting valid but rare modes, limiting the discovery of novel solutions. Existing approaches, such as guidance methods or combining reinforcement learning with diffusion, either push samples into infeasible regions or struggle to escape local minima, failing to systematically uncover diverse behaviors. To address these challenges, we propose a framework that combines Feynman-Kac correctors with a novel guiding potential that systematically guides diffusion policy samples towards promising yet underrepresented samples. These trajectories are refined using sampling-based trajectory optimization and reincorporated into the training set to retrain the diffusion policy. Our method effectively mines and repairs novel trajectories, enabling the systematic discovery of diverse and executable behaviors. We demonstrate the effectiveness of our framework across a range of manipulation environments, consistently discovering new behaviors.
- Abstract(参考訳): 拡散モデルは、多モーダルな行動-軌道分布のモデル化に優れた拡散ポリシーを持つ、ロボット工学における生成モデリングの強力なツールとなっている。
しかしながら、デモが限定されている場合、標準サンプリングはしばしば、有効だが稀なモードを無視しながら支配的な振る舞いを再現し、新しい解の発見を制限する。
指導方法や強化学習と拡散を組み合わせた既存のアプローチは、サンプルを実用不可能な地域にプッシュするか、あるいは局所的なミニマから逃れるのに苦労し、体系的に多様な行動を明らかにするのに失敗する。
これらの課題に対処するため,Feynman-Kac 補正器と拡散政策サンプルを体系的に案内する新たな誘導電位を組み合わせたフレームワークを提案する。
これらのトラジェクトリはサンプリングベースのトラジェクトリ最適化を用いて洗練され、拡散政策を再訓練するためのトレーニングセットに再組み込まれる。
提案手法は, 新規な軌道を効果的にマイニングし, 修復し, 多様な, 実行可能な動作の体系的な発見を可能にする。
我々は様々な操作環境にまたがってフレームワークの有効性を実証し、新しい振る舞いを継続的に発見する。
関連論文リスト
- Scalable Discrete Diffusion Samplers: Combinatorial Optimization and Statistical Physics [7.873510219469276]
離散拡散サンプリングのための2つの新しいトレーニング手法を提案する。
これらの手法は、メモリ効率のトレーニングを行い、教師なし最適化の最先端結果を達成する。
SN-NISとニューラルチェインモンテカルロの適応を導入し,離散拡散モデルの適用を初めて可能とした。
論文 参考訳(メタデータ) (2025-02-12T18:59:55Z) - Adaptive teachers for amortized samplers [76.88721198565861]
そこで,本研究では,初等無罪化標本作成者(学生)の指導を指導する適応的学習分布(教師)を提案する。
本研究では, この手法の有効性を, 探索課題の提示を目的とした合成環境において検証する。
論文 参考訳(メタデータ) (2024-10-02T11:33:13Z) - GUIDE: Guidance-based Incremental Learning with Diffusion Models [3.046689922445082]
GUIDEは,拡散モデルからサンプルのリハーサルを誘導する,新しい連続学習手法である。
実験の結果,GUIDEは破滅的忘れを著しく減らし,従来のランダムサンプリング手法より優れ,生成的再生を伴う継続的な学習における最近の最先端の手法を超越した。
論文 参考訳(メタデータ) (2024-03-06T18:47:32Z) - Improved off-policy training of diffusion samplers [93.66433483772055]
本研究では,非正規化密度やエネルギー関数を持つ分布からサンプルを抽出する拡散モデルの訓練問題について検討する。
シミュレーションに基づく変分法や非政治手法など,拡散構造推論手法のベンチマークを行った。
我々の結果は、過去の研究の主張に疑問を投げかけながら、既存のアルゴリズムの相対的な利点を浮き彫りにした。
論文 参考訳(メタデータ) (2024-02-07T18:51:49Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。