論文の概要: CoJEPA: Combining Contrastive Learning and JEPA for Global-Local Music Representations
- arxiv url: http://arxiv.org/abs/2608.30974v2
- Date: Sun, 06 Sep 2026 08:38:39 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-09 22:37:20.324156
- Title: CoJEPA: Combining Contrastive Learning and JEPA for Global-Local Music Representations
- Title(参考訳): CoJEPA:グローバルローカル音楽表現のためのコントラスト学習とJEPAを組み合わせる
- Authors: Gabriel Meseguer-Brocal, Yuexuan Kong, Romain Hennequin,
- Abstract要約: JEPA(Joint-Embedding Predictive Architecture)は、潜在空間における自己教師付き予測を通じて、豊かな表現を学習する上で、強力なパフォーマンスを示している。
コントラスト学習は、訓練に安定し、強力なグローバル表現を生み出すが、その目的のグローバルな性質によって局所的なタスクに制限される。
マスクされたシーケンストークンでJEPAオブジェクトと共同でトレーニングされた単一の共有バックボーンと、クラストークンで対照的な目的である。
- 参考スコア(独自算出の注目度): 12.163372992542326
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: Joint-Embedding Predictive Architecture (JEPA) has shown strong performance in learning rich representations through self-supervised prediction in latent space. However, it typically relies on teacher--student architecture with an EMA to stabilise training, and can tend to yield uninformative representations. Contrastive learning is stable to train and produces strong global representations, but remains limited on local tasks by the global nature of its objective. In this work, we combine both into CoJEPA: a single shared backbone jointly trained with a JEPA objective on masked sequence tokens and a contrastive objective on the class token. The contrastive gradient provides stability, removing the need for an EMA teacher entirely, while JEPA enriches the sequence tokens via local predictions that contrastive learning alone cannot provide. Crucially, no extra parameters are added to the backbone: the same model is guided towards richer representations purely through the design of its training signal. CoJEPA takes the best of both worlds, outperforming or matching both individual methods across global and local MIR tasks, with a particularly strong advantage on tonal and harmonic understanding, and without any task-specific architectural changes. CoJEPA shows that combining objectives with complementary inductive biases can substitute for scale, encouraging future work to invest in smarter training objectives over ever-larger models.
- Abstract(参考訳): JEPA(Joint-Embedding Predictive Architecture)は、潜在空間における自己教師付き予測を通じて、豊かな表現を学習する上で、強力なパフォーマンスを示している。
しかし、通常は教師-学生アーキテクチャに頼り、EMAを使ってトレーニングを安定化し、非形式的な表現を産み出す傾向がある。
コントラスト学習は、訓練に安定し、強力なグローバル表現を生み出すが、その目的のグローバルな性質によって局所的なタスクに制限される。
マスクされたシーケンストークンでJEPAオブジェクトと共同でトレーニングされた単一の共有バックボーンと、クラストークンで対照的な目的である。
対照的な勾配は安定性を提供し、EMA教師の必要性を完全に取り除き、JEPAは対照的な学習だけでは提供できないローカル予測を通じてシーケンストークンを豊かにする。
重要なことに、バックボーンに余分なパラメータを追加することはない。同じモデルは、トレーニング信号の設計を通じて、純粋によりリッチな表現へと導かれる。
CoJEPAは、グローバルなMIRタスクとローカルなMIRタスクにまたがって、個々のメソッドをパフォーマンスまたはマッチングすることで、両方の世界の長所を取ります。
CoJEPAは、目的と相補的な帰納的バイアスを組み合わせることは、スケールに取って代わる可能性があり、今後はより大規模なモデルよりもスマートなトレーニング目標に投資することを奨励している。
関連論文リスト
- Dreaming Of Others: Latent Teammate Modeling In World Models For Multi-Agent Reinforcement Learning [0.0]
本稿では,Dreamerスタイルのリカレントステートスペースモデルの潜在状態を,環境やチームメイトコンポーネントに分解するアーキテクチャを提案する。
これらのチームメイトは俳優と批評家を条件付けし、エージェントが様々な協力者に想像し適応できるようにする。
論文 参考訳(メタデータ) (2026-05-29T14:34:50Z) - UWM-JEPA: Predictive World Models That Imagine in Belief Space [0.2864713389096699]
本稿では,JEPAの世界モデルであるUnitary World Model JEPAを紹介した。
この構造はロールアウト中に関節状態スペクトルを正確に保存するため、予測器自体が表現された不確かさを解消することはできない。
JEPAの世界モデルでは、部分的な可観測性、潜伏幾何学、予測力学が重要であり、フリーズされたコンテキストエンコーディング能力だけではありません。
論文 参考訳(メタデータ) (2026-05-25T00:28:51Z) - LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics [53.247652209132376]
JEPA(Joint-Embedding Predictive Architectures)は、有望な青写真を提供するが、実践的なガイダンスや理論の欠如がアドホックな研究開発につながっている。
我々はJEPAの包括的な理論を示し、それをbf LeJEPAでインスタンス化する。
論文 参考訳(メタデータ) (2025-11-11T18:21:55Z) - ACT-JEPA: Novel Joint-Embedding Predictive Architecture for Efficient Policy Representation Learning [90.41852663775086]
ACT-JEPAは模倣学習と自己教師型学習を統合する新しいアーキテクチャである。
我々はアクションシーケンスと抽象的な観察シーケンスを予測するポリシーを訓練する。
実験の結果,ACT-JEPAは時間環境の動的学習によって表現の質を向上させることがわかった。
論文 参考訳(メタデータ) (2025-01-24T16:41:41Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。