論文の概要: FLUID: Continuous-Time Hyperconnected Sparse Transformer for Sink-Free Learning
- arxiv url: http://arxiv.org/abs/2605.04421v1
- Date: Wed, 06 May 2026 02:27:25 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-05-07 18:41:07.609963
- Title: FLUID: Continuous-Time Hyperconnected Sparse Transformer for Sink-Free Learning
- Title(参考訳): FLUID: シンクフリー学習のための連続時間ハイパーコネクテッドスパース変換器
- Authors: Waleed Razzaq, Yun-Bo Zhao,
- Abstract要約: 連続時間(CT)変換器は、連続力学を用いた入力や出力埋め込みを利用して、CT-RNN上の不規則および長距離モデリングを改善する。
我々は、連続力学を直接注意計算に組み込むCT変換器FLUIDを提案し、それをLiquid Attention Network(LAN)に置き換える。
本研究では, (i) 不規則な時系列, (ii) 長距離モデリング, (iii) 自動運転車の車線維持制御, (iv) 不足データ体制下での物理力学の学習など,幅広い学習課題についてFLUIDを評価した。
- 参考スコア(独自算出の注目度): 2.312232949770907
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Continuous-time (CT) Transformers improve irregular and long-range modeling over CT-RNNs by exploiting inputs or outputs embeddings with continuous dynamics. However, the core scaled-dot-product-attention (SDPA) mechanism remains inherently discrete. We propose FLUID (Flexible Unified Information Dynamics), a CT Transformer that incorporates continuous dynamics directly into the attention computation by replacing it with Liquid Attention Network (LAN). LAN reinterprets attention logits as continuous dynamical system and reformulates them as the solution to a linear ODE modulated by input-dependent nonlinear recurrent gates. Theoretically, we establish stability guarantees for LAN dynamics and show that it serves as an interpolating middle ground between SDPA and CT-RNNs, recovering each as special case under well-defined parameterization of its gating functions. LAN also introduces an explicit attention-sink gate to eliminate disproportionate attention mass on uninformative nodes. FLUID replaces standard residual connections with input-dependent Liquid Hyper-Connections to adaptively regulate interlayer information flow. Empirically, we evaluate FLUID on a broad set of learning tasks, including (i) irregular time-series, (ii) long-range modeling, (iii) lane-keeping control of autonomous vehicles, and (iv) learning physical dynamics under a scarce data regime. Across all the tasks, FLUID consistently matches or outperforms CT baselines, achieving improvements of up to 47% in certain scenarios and enhancing generalization under distributional shifts. Additionally, FLUID demonstrates superior noise robustness and a self-correcting inductive bias in autonomous vehicle control. We also provide a detailed analysis of key hyperparameters to guide tuning and show that FLUID occupies an intermediate position among competing approaches in terms of runtime and memory efficiency.
- Abstract(参考訳): 連続時間(CT)変換器は、連続力学を用いた入力や出力埋め込みを利用して、CT-RNN上の不規則および長距離モデリングを改善する。
しかし、コアスケールド・ドット・プロダクツ・アテンション(SDPA)機構は本質的には離散的である。
FLUID(Flexible Unified Information Dynamics, フレキシブル・ユニファイド・インフォメーション・ダイナミクス)は, 連続力学を直接注意計算に組み込むCT変換器である。
LANはアテンションロジットを連続力学系として再解釈し、入力依存の非線形リカレントゲートによって変調された線形ODEの解として再構成する。
理論的には、LAN力学の安定性保証を確立し、SDPAとCT-RNNの間を補間する中間地盤として機能し、それぞれがゲーティング関数のパラメータ化を適切に定義した特別ケースとして回復することを示す。
LANはまた、不整形ノードに対する不均等な注意質量を排除するために、明示的なアテンションシンクゲートも導入している。
FLUIDは、標準の残コネクションを入力依存のLiquid Hyper-Connectionsに置き換えて、層間情報の流れを適応的に制御する。
経験的に、我々はFLUIDを幅広い学習課題において評価する。
(i)不規則な時間帯
(II)長距離モデリング
三 自動運転車の車線維持制御及び
(4)少ないデータ体制下で物理力学を学ぶこと。
すべてのタスクにおいて、FLUIDは一貫してCTベースラインに適合し、特定のシナリオで最大47%の改善を実現し、分散シフト下での一般化を向上する。
さらに、FLUIDは、自律走行車の制御において、優れたノイズ堅牢性と自己補正誘導バイアスを示す。
また,FLUIDが実行時とメモリ効率の面で競合するアプローチの中間的な位置を占めることを示すために,鍵パラメータの詳細な解析を行った。
関連論文リスト
- Drift-Aware Online Dynamic Learning for Nonstationary Multivariate Time Series: Application to Sintering Quality Prediction [8.77587871192859]
堅牢な多時間予測性能を維持するために,DA-MSDL(Drift-Aware Multi-Scale Dynamic Learning)フレームワークを提案する。
提案フレームワークは,非定常環境における品質モニタリングに有効なオンライン動的学習パラダイムを提供する。
論文 参考訳(メタデータ) (2026-04-10T14:30:57Z) - Extended Hybrid Timed Petri Nets with Semi-Supervised Anomaly Detection for Switched Systems, Modelling and Fault Detection [0.0]
本稿では,ハイブリッド力学系のための統合障害検出フレームワークを提案する。
拡張時間連続ペトリネットモデルと半教師付き異常検出を統合している。
その結果,高い検出精度,高速収束,ロバスト性能が得られた。
論文 参考訳(メタデータ) (2026-04-05T10:54:15Z) - Real-Time Generative Policy via Langevin-Guided Flow Matching for Autonomous Driving [22.3805998088591]
DACER-Fは、自律運転システムにおける生成ポリシーのフローマッチングアルゴリズムである。
ヒューマノイド・スタンド・タスクで775.8のスコアを獲得し、以前の手法を上回ります。
論文 参考訳(メタデータ) (2026-03-03T05:35:53Z) - Causal Autoregressive Diffusion Language Model [70.7353007255797]
CARDは厳密な因果注意マスク内の拡散過程を再構成し、単一の前方通過で密集した1対1の監視を可能にする。
我々の結果は,CARDが並列生成のレイテンシの利点を解放しつつ,ARMレベルのデータ効率を実現することを示す。
論文 参考訳(メタデータ) (2026-01-29T17:38:29Z) - FAIM: Frequency-Aware Interactive Mamba for Time Series Classification [87.84511960413715]
時系列分類(TSC)は、環境モニタリング、診断、姿勢認識など、多くの実世界の応用において重要である。
本稿では,周波数対応対話型マンバモデルであるFAIMを提案する。
FAIMは既存の最先端(SOTA)手法を一貫して上回り、精度と効率のトレードオフが優れていることを示す。
論文 参考訳(メタデータ) (2025-11-26T08:36:33Z) - Optimal Control Meets Flow Matching: A Principled Route to Multi-Subject Fidelity [35.95129874095729]
テキスト・トゥ・イメージ(T2I)モデルは単一エンタリティ・プロンプトに優れるが、多目的記述に苦慮する。
マルチオブジェクト忠実度に向けてサンプリングダイナミクスを操るための原理的最適化可能な目的を持った最初の理論的枠組みを導入する。
論文 参考訳(メタデータ) (2025-10-02T17:59:58Z) - Vaccinating Federated Learning for Robust Modulation Classification in Distributed Wireless Networks [0.0]
雑音レベルの異なる信号間の一般化性向上を目的とした新しいAMCモデルであるFedVaccineを提案する。
FedVaccineは、分割学習戦略を用いることで、既存のFLベースのAMCモデルの線形集約の限界を克服する。
これらの結果は、実用的な無線ネットワーク環境におけるAMCシステムの信頼性と性能を高めるためのFedVaccineの可能性を浮き彫りにした。
論文 参考訳(メタデータ) (2024-10-16T17:48:47Z) - Leveraging Low-Rank and Sparse Recurrent Connectivity for Robust
Closed-Loop Control [63.310780486820796]
繰り返し接続のパラメータ化が閉ループ設定のロバスト性にどのように影響するかを示す。
パラメータが少ないクローズドフォーム連続時間ニューラルネットワーク(CfCs)は、フルランクで完全に接続されたニューラルネットワークよりも優れています。
論文 参考訳(メタデータ) (2023-10-05T21:44:18Z) - Towards Long-Term Time-Series Forecasting: Feature, Pattern, and
Distribution [57.71199089609161]
長期的時系列予測(LTTF)は、風力発電計画など、多くのアプリケーションで需要が高まっている。
トランスフォーマーモデルは、高い計算自己認識機構のため、高い予測能力を提供するために採用されている。
LTTFの既存の手法を3つの面で区別する,Conformer という,効率的なTransformer ベースモデルを提案する。
論文 参考訳(メタデータ) (2023-01-05T13:59:29Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。