論文の概要: FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models
- arxiv url: http://arxiv.org/abs/2607.24522v1
- Date: Mon, 27 Jul 2026 15:03:22 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-28 22:34:15.472946
- Title: FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models
- Title(参考訳): FlowCTS: フローモデルのオンライン連続軌道推定
- Abstract要約: フロー連続軌道スーパービジョン(FlowCTS)を提案する。
FlowCTSは、同じ学生訪問状態の学生と参照軌跡とを一致させる。
マルチ参照設定では、単一状態のFlowCTS-OPDはバニラKLベースのOPDよりも高速な収束で性能が向上する。
- 参考スコア(独自算出の注目度): 65.32587991232592
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: While on-policy distillation (OPD) effectively addresses sparse rewards and exposure bias in large language model post-training, its extension to flow models remains underexplored. To this end, we propose Flow Continuous Trajectory Supervision (FlowCTS), which matches subsequent student and reference trajectories initialized from the same student-visited state. Using the integral relation between trajectories and velocity fields, we derive a temporally weighted velocity-matching upper bound and discretize it into practical objectives parameterized by the number of supervision steps. Under a multi-reference setup, single-state FlowCTS-OPD outperforms vanilla KL-based OPD with faster convergence. FlowCTS-OPD improves GenEval from 0.90 to 0.93, OCR from 0.90 to 0.92, and PickScore from 22.75 to 23.06, while outperforming a mixed-reward RL baseline across all target metrics. Further analysis reveals a clear temporal supervision mismatch in vanilla KL-based OPD arising from its auxiliary SDE transition kernels. Beyond on-policy setting,FlowCTS also consistently outperforms vanilla SFT , particularly on OCR, while increasing supervision steps exhibit a trade-off between richer trajectory information and greater optimization difficulty.
- Abstract(参考訳): オンライン蒸留(OPD)は、大規模言語モデルの訓練後におけるスパース報酬と露出バイアスに効果的に対処するが、フローモデルへの拡張は未解明のままである。
この目的のために,フロー連続軌道スーパービジョン (FlowCTS) を提案する。
軌道と速度場の積分関係を用いて、時間重み付けされた速度マッチング上界を導出し、監督ステップの数によってパラメータ化された現実的な目的に識別する。
マルチ参照設定では、単一状態のFlowCTS-OPDはバニラKLベースのOPDよりも高速な収束で性能が向上する。
FlowCTS-OPD は GenEval を 0.90 から 0.93 に改善し、OCR は 0.90 から 0.92 に、PickScore は 22.75 から 23.06 に改善した。
さらなる分析により、補助的なSDE遷移カーネルから生じるバニラKLベースのOPDの時間的監督ミスマッチが明らかになる。
オンライン設定以外にも、FlowCTSはバニラSFT、特にOCRよりも一貫して優れており、監督手順の増大は、よりリッチな軌道情報と最適化の難しさの間のトレードオフを示している。
関連論文リスト
- DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models [55.01951088768769]
DiffusionOPDはオンライン政策蒸留(OPD)に基づく拡散モデルのための新しいマルチタスクトレーニングパラダイムである
本研究では,DiffusionOPDがトレーニング効率と最終性能において,マルチリワードRLとカスケードRLのベースラインを一貫して上回っていることを示す。
論文 参考訳(メタデータ) (2026-05-14T16:49:09Z) - Preserving Foundational Capabilities in Flow-Matching VLAs through Conservative SFT [12.18095306054314]
ConSFT(Reserve Supervised Fine-Tuning)は、壊滅的な忘れを軽減しつつ、ターゲット分布に適応する最適化目標である。
LIBERO と RoboTwin ベンチマークの ConSFT を最先端のフローマッチング VLA で評価した。
論文 参考訳(メタデータ) (2026-05-09T10:59:03Z) - EAGLE: Edge-Aware Graph Learning for Proactive Delivery Delay Prediction in Smart Logistics Networks [0.5653106385738822]
本稿では,プロアクティブサプライチェーンリスク管理のためのハイブリッドディープラーニングフレームワークを提案する。
提案手法は,軽量な Transformer patch encoder を用いて時間次数フローのダイナミクスを共同でモデル化する。
このフレームワークは0.0089のクロスシードF1標準偏差を示しており、最高の改良版よりも3.8倍改善されている。
論文 参考訳(メタデータ) (2026-04-06T23:35:15Z) - On the Role of Reasoning Patterns in the Generalization Discrepancy of Long Chain-of-Thought Supervised Fine-Tuning [63.41902113656453]
長いチェーン・オブ・ソート(CoT)軌道上でのSFT(Supervised Fine-Tuning)は、大きな推論モデルを構築する上で重要なフェーズとなっている。
2つの競合モデルによって生成された2つの検証されたCoT軌道源を用いて比較研究を行う。
textttDeepSeek-R1-0528データ上のSFTは、トレーニング損失を著しく低減するが、一般化性能は著しく低下する。
論文 参考訳(メタデータ) (2026-04-02T07:00:54Z) - Temporal Pair Consistency for Variance-Reduced Flow Matching [13.328987133593154]
TPC(Temporal Pair Consistency)は、同じ確率経路に沿ってペア化された時間ステップで速度予測を結合する軽量な分散還元原理である。
フローマッチング内で確立されたTPCは、複数の解像度でCIFAR-10とImageNetのサンプル品質と効率を改善する。
論文 参考訳(メタデータ) (2026-02-04T00:05:21Z) - DiffusionNFT: Online Diffusion Reinforcement with Forward Process [99.94852379720153]
Diffusion Negative-aware FineTuning (DiffusionNFT) は、フローマッチングを通じて前方プロセス上で直接拡散モデルを最適化する新しいオンラインRLパラダイムである。
DiffusionNFTは、CFGフリーのFlowGRPOよりも25倍効率が高い。
論文 参考訳(メタデータ) (2025-09-19T16:09:33Z) - Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections [65.36449542323277]
本稿では,Large Language Model (LLM) 後の学習において,SFT(Supervised Fine-Tuning) と優先学習を統合した理論フレームワークを提案する。
そこで本研究では,学習率の簡易かつ効果的な削減手法を提案する。
論文 参考訳(メタデータ) (2025-06-15T05:42:29Z) - CT-OT Flow: Estimating Continuous-Time Dynamics from Discrete Temporal Snapshots [14.39808568320318]
Continuous-Time Optimal Transport Flow (CT-OT Flow)は、高解像度のタイムラベルを推論し、連続時間データ分散を再構築するフレームワークである。
CT-OT Flow は OT-CFM, [SF](2)M, TrajectoryNet, MFM, ENOT と比較して分布誤差と軌道誤差を低減する。
論文 参考訳(メタデータ) (2025-05-23T00:12:49Z) - FlowTS: Time Series Generation via Rectified Flow [67.41208519939626]
FlowTSは、確率空間における直線輸送を伴う整流フローを利用するODEベースのモデルである。
非条件設定では、FlowTSは最先端のパフォーマンスを達成し、コンテキストFIDスコアはStockとETThデータセットで0.019と0.011である。
条件設定では、太陽予測において優れた性能を達成している。
論文 参考訳(メタデータ) (2024-11-12T03:03:23Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。