論文の概要: Towards Multimodal Time Series Anomaly Detection with Semantic Alignment and Condensed Interaction
- arxiv url: http://arxiv.org/abs/2603.21612v1
- Date: Mon, 23 Mar 2026 06:18:23 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-03-24 19:11:39.520484
- Title: Towards Multimodal Time Series Anomaly Detection with Semantic Alignment and Condensed Interaction
- Title(参考訳): セマンティックアライメントと凝縮相互作用を用いたマルチモーダル時系列異常検出に向けて
- Authors: Shiyan Hu, Jianxin Jin, Yang Shu, Peng Chen, Bin Yang, Chenjuan Guo,
- Abstract要約: 時系列異常検出は多くの力学系において重要な役割を果たす。
従来の手法は主に単調な数値データに依存していた。
我々は,新しいマルチモーダル時系列異常検出モデル(MindTS)を提案する。
- 参考スコア(独自算出の注目度): 20.045235276119595
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Time series anomaly detection plays a critical role in many dynamic systems. Despite its importance, previous approaches have primarily relied on unimodal numerical data, overlooking the importance of complementary information from other modalities. In this paper, we propose a novel multimodal time series anomaly detection model (MindTS) that focuses on addressing two key challenges: (1) how to achieve semantically consistent alignment across heterogeneous multimodal data, and (2) how to filter out redundant modality information to enhance cross-modal interaction effectively. To address the first challenge, we propose Fine-grained Time-text Semantic Alignment. It integrates exogenous and endogenous text information through cross-view text fusion and a multimodal alignment mechanism, achieving semantically consistent alignment between time and text modalities. For the second challenge, we introduce Content Condenser Reconstruction, which filters redundant information within the aligned text modality and performs cross-modal reconstruction to enable interaction. Extensive experiments on six real-world multimodal datasets demonstrate that the proposed MindTS achieves competitive or superior results compared to existing methods. The code is available at: https://github.com/decisionintelligence/MindTS.
- Abstract(参考訳): 時系列異常検出は多くの力学系において重要な役割を果たす。
その重要性にも拘わらず、従来のアプローチは、他のモダリティからの相補的な情報の重要性を見越して、主に単調な数値データに依存してきた。
本稿では,(1)異種マルチモーダルデータ間のセマンティックな整合性を実現する方法,(2)冗長なモーダル情報をフィルタリングして相互モーダル間相互作用を効果的に強化する方法の2つの課題に対処することに焦点を当てた,新しいマルチモーダル時系列異常検出モデル(MindTS)を提案する。
最初の課題に対処するために、細粒度タイムテキストセマンティックアライメントを提案する。
クロスビューテキスト融合とマルチモーダルアライメント機構を通じて、外因性および内因性テキスト情報を統合し、時間とテキストのモダリティを意味的に一貫したアライメントを実現する。
第2の課題としてコンテント・コンデンサ・リコンストラクション(Content Condenser Reコンストラクション)を提案する。
6つの実世界のマルチモーダルデータセットに関する大規模な実験は、提案されたMindTSが既存の手法と比較して競争力または優れた結果が得られることを示した。
コードはhttps://github.com/decisionintelligence/MindTSで公開されている。
関連論文リスト
- FindRec: Stein-Guided Entropic Flow for Multi-Modal Sequential Recommendation [57.577843653775]
textbfFindRec (textbfFlexible unified textbfinformation textbfdisentanglement for multi-modal sequence textbfRecommendation)を提案する。
Stein kernel-based Integrated Information Coordination Module (IICM) は理論上、マルチモーダル特徴とIDストリーム間の分散一貫性を保証する。
マルチモーダル特徴を文脈的関連性に基づいて適応的にフィルタリング・結合するクロスモーダル・エキスパート・ルーティング機構。
論文 参考訳(メタデータ) (2025-07-07T04:09:45Z) - Multi-modal Time Series Analysis: A Tutorial and Survey [36.93906365779472]
マルチモーダル時系列分析はデータマイニングにおいて顕著な研究領域となっている。
しかし、マルチモーダル時系列の効果的な解析は、データの不均一性、モダリティギャップ、不整合、固有ノイズによって妨げられる。
マルチモーダル時系列法の最近の進歩は、クロスモーダル相互作用を通じて、マルチモーダルコンテキストを利用した。
論文 参考訳(メタデータ) (2025-03-17T20:30:02Z) - Joint Fusion and Encoding: Advancing Multimodal Retrieval from the Ground Up [26.32353412029717]
今日のインターネットアプリケーションには情報検索が不可欠である。
伝統的なセマンティックマッチング技術は、細粒なクロスモーダル相互作用を捉えるのにしばしば不足する。
我々は、視覚的およびテキスト的手がかりをゼロから融合する統合検索フレームワークを導入する。
論文 参考訳(メタデータ) (2025-02-27T11:41:55Z) - Detecting Misinformation in Multimedia Content through Cross-Modal Entity Consistency: A Dual Learning Approach [10.376378437321437]
クロスモーダルなエンティティの整合性を利用して、ビデオコンテンツから誤情報を検出するためのマルチメディア誤情報検出フレームワークを提案する。
以上の結果から,MultiMDは最先端のベースラインモデルより優れていることが示された。
論文 参考訳(メタデータ) (2024-08-16T16:14:36Z) - Asynchronous Multimodal Video Sequence Fusion via Learning Modality-Exclusive and -Agnostic Representations [19.731611716111566]
本稿では,モダリティ学習のためのマルチモーダル融合手法を提案する。
我々は、モーダル内の信頼性のあるコンテキストダイナミクスをキャプチャする予測的自己アテンションモジュールを導入する。
階層的クロスモーダルアテンションモジュールは、モダリティ間の価値ある要素相関を探索するために設計されている。
両識別器戦略が提示され、異なる表現を敵対的に生成することを保証する。
論文 参考訳(メタデータ) (2024-07-06T04:36:48Z) - From Text to Pixels: A Context-Aware Semantic Synergy Solution for
Infrared and Visible Image Fusion [66.33467192279514]
我々は、テキスト記述から高レベルなセマンティクスを活用し、赤外線と可視画像のセマンティクスを統合するテキスト誘導多モード画像融合法を提案する。
本手法は,視覚的に優れた融合結果を生成するだけでなく,既存の手法よりも高い検出mAPを達成し,最先端の結果を得る。
論文 参考訳(メタデータ) (2023-12-31T08:13:47Z) - Multi-Grained Multimodal Interaction Network for Entity Linking [65.30260033700338]
マルチモーダルエンティティリンクタスクは、マルチモーダル知識グラフへの曖昧な言及を解決することを目的としている。
MELタスクを解決するための新しいMulti-Grained Multimodal InteraCtion Network $textbf(MIMIC)$ frameworkを提案する。
論文 参考訳(メタデータ) (2023-07-19T02:11:19Z) - Align and Attend: Multimodal Summarization with Dual Contrastive Losses [57.83012574678091]
マルチモーダル要約の目標は、異なるモーダルから最も重要な情報を抽出し、出力要約を形成することである。
既存の手法では、異なるモダリティ間の時間的対応の活用に失敗し、異なるサンプル間の本質的な相関を無視する。
A2Summ(Align and Attend Multimodal Summarization)は、マルチモーダル入力を効果的に整列し、参加できる統一型マルチモーダルトランスフォーマーモデルである。
論文 参考訳(メタデータ) (2023-03-13T17:01:42Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。