論文の概要: Douyin Multimodal Embedding Model Technical Report
- arxiv url: http://arxiv.org/abs/2608.02148v1
- Date: Mon, 03 Aug 2026 12:31:49 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-04 15:07:25.548671
- Title: Douyin Multimodal Embedding Model Technical Report
- Title(参考訳): Douyin Multimodal Embedding Model 技術報告
- Abstract要約: マルチモーダル表現学習は、現代のAIの基盤となっている。
両強度を組み合わせた2段階のモデルであるDME(Douyin Multimodal Embedding)を提案する。
- 参考スコア(独自算出の注目度): 77.9018834072698
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vectors, it powers industrial search and recommendation and underpins modern agents. Real-world platforms with complex modalities and massive-scale content, such as Douyin, Xiaohongshu, and YouTube, demand both efficiency under billion-scale indexing and fine-grained discrimination for hard matching. Existing MLLM embedding models rarely satisfy both. Contrastive models are efficient but rely on pair-level supervision too coarse for fine-grained distinctions, while CoT-based models improve discrimination through explicit generation impractical to serve online. We present Douyin Multimodal Embedding (DME), a model trained in two stages to combine both strengths. Stage 1 performs large-scale contrastive pre-training that establishes a unified multimodal embedding space with broad modality and task coverage. Stage 2 supplements semantic sufficiency, the property that an embedding is grounded in retrieval-relevant evidence and preserves fine-grained counterpart-side semantics, via two mechanisms. Evidence-Grounded Typed Latent Reasoning organizes retrieval evidence through hidden-space latent reasoning, and Cross-Conditional Reconstruction enforces counterpart-side semantics through cross-directional autoregressive reconstruction. Both act only during training and add only marginal query-side overhead, so DME serves as efficiently as a standard contrastive encoder. On MMEB-v2, DME reaches state-of-the-art results at comparable scales for its 2B and 9B variants (74.8 and 78.4), with especially strong video and visual-document tasks. In production, DME delivers a 2.92% relative gain on Douyin's in-house offline evaluation set, is deployed across Douyin scenarios such as generative, image, and AI search, and yields a 0.1% Lifetime (LT) gain in online A/B testing on Douyin search.
- Abstract(参考訳): マルチモーダル表現学習は現代のAIの基礎である。
マルチモーダルクエリとターゲットをベクトルに符号化することで、産業検索とレコメンデーションの力となり、現代のエージェントを支える。
Douyin、Xiaohongshu、YouTubeのような複雑なモダリティと大規模なコンテンツを持つ現実世界のプラットフォームは、数十億ドル規模のインデックス付けの下での効率と、ハードマッチングのためのきめ細かい識別の両方を要求する。
既存のMLLM埋め込みモデルは両方を満たすことは滅多にない。
コントラストモデルは効率的だが、ペアレベルの監督は微妙な区別をしすぎるが、CoTベースのモデルはオンラインサービスを提供するための明示的な生成非現実性を通じて差別を改善する。
両強度を組み合わせた2段階のモデルであるDME(Douyin Multimodal Embedding)を提案する。
ステージ1は、広範囲なモダリティとタスクカバレッジを備えた統合マルチモーダル埋め込み空間を確立する大規模なコントラスト事前訓練を行う。
ステージ2はセマンティック・セマンティクスを補うもので、埋め込みは検索関連エビデンスに基礎を置いており、2つのメカニズムを通じてきめ細かい相補的なセマンティクスを保存している。
Evidence-Grounded Typed Latent Reasoningは隠れ空間潜在推論を通じて検索証拠を整理し、Cross-Conditional Reasoningは、双方向の自己回帰的再構築を通じて相互のセマンティクスを強制する。
どちらもトレーニング時にのみ動作し、限界クエリ側のオーバーヘッドのみを追加するため、DMEは標準のコントラストエンコーダとして効率的に機能する。
MMEB-v2では、DMEは2Bと9Bの変種(74.8と78.4)に匹敵する規模で最先端の成果を収めている。
実運用では、DMEはDouyinの社内オフライン評価セットに対して2.92%の相対的な利得を提供し、生成、画像、AI検索といったDouyinのシナリオにまたがってデプロイされ、オンラインA/B検索では0.1%のライフタイム(LT)の利得が得られる。
関連論文リスト
- VTFusion: A Vision-Text Multimodal Fusion Network for Few-Shot Anomaly Detection [24.88767599022225]
Few-Shot Anomaly Detection (FSAD) は、希少な正規参照を用いて不規則を識別するための重要なパラダイムとして登場した。
本研究では,FSADに適した視覚テキスト多モード融合フレームワークであるVTFusionを提案する。
論文 参考訳(メタデータ) (2026-01-23T00:30:24Z) - SERM: Self-Evolving Relevance Model with Agent-Driven Learning from Massive Query Streams [53.78257200138774]
本稿では,2つの相補的マルチエージェントモジュールからなる自己進化関連モデル(SERM)を提案する。
我々はSERMを大規模産業環境で評価し、毎日数十億のユーザリクエストを処理している。
論文 参考訳(メタデータ) (2026-01-14T14:31:16Z) - Filling the Gaps: A Multitask Hybrid Multiscale Generative Framework for Missing Modality in Remote Sensing Semantic Segmentation [28.992992584085787]
マルチモーダル学習は、通常の単調モデルと比較して大きな性能向上を示した。
現実のシナリオでは、センサーの故障と悪天候のためにマルチモーダル信号が欠落する可能性がある。
本稿では,これらの制約に対処するために,GEMMNet(Generative-Enhanced MultiModal Learning Network)を提案する。
論文 参考訳(メタデータ) (2025-09-14T05:40:35Z) - Progressive Semantic Residual Quantization for Multimodal-Joint Interest Modeling in Music Recommendation [6.790539226766362]
本稿では,2段階の新たなマルチモーダルレコメンデーションフレームワークを提案する。
最初の段階では、モーダル固有およびモーダルジョイントのセマンティックIDを生成する。
第2段階では、ユーザのマルチモーダルな関心をモデル化するために、マルチコードブックのクロスアテンションネットワークが設計されている。
論文 参考訳(メタデータ) (2025-08-28T02:16:57Z) - MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings [75.0617088717528]
MoCaは、トレーニング済みのVLMバックボーンを効果的な双方向埋め込みモデルに変換するためのフレームワークである。
MoCaは、MMEBとViDoRe-v2ベンチマークのパフォーマンスを継続的に改善し、新しい最先端の結果を達成する。
論文 参考訳(メタデータ) (2025-06-29T06:41:00Z) - Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark [69.02666229531322]
モダリティ不完全産業異常検出(MIIAD)の先駆的研究を紹介する。
その結果,既存のMIAD手法はMIIADベンチでは性能が悪く,性能が著しく低下していることが判明した。
本稿では,新しい2段階のロバストモードアリティファジングと検出フレームwoRk(RADAR)を提案する。
論文 参考訳(メタデータ) (2024-10-02T16:47:55Z) - PMT: Progressive Mean Teacher via Exploring Temporal Consistency for Semi-Supervised Medical Image Segmentation [51.509573838103854]
医用画像セグメンテーションのための半教師付き学習フレームワークであるプログレッシブ平均教師(PMT)を提案する。
我々のPMTは、トレーニングプロセスにおいて、堅牢で多様な特徴を学習することで、高忠実な擬似ラベルを生成する。
CT と MRI の異なる2つのデータセットに対する実験結果から,本手法が最先端の医用画像分割法より優れていることが示された。
論文 参考訳(メタデータ) (2024-09-08T15:02:25Z) - Efficient Multimodal Transformer with Dual-Level Feature Restoration for
Robust Multimodal Sentiment Analysis [47.29528724322795]
マルチモーダルセンシング分析(MSA)が近年注目を集めている。
著しい進歩にもかかわらず、堅牢なMSAへの道にはまだ2つの大きな課題がある。
デュアルレベル特徴回復 (EMT-DLFR) を用いた高効率マルチモーダル変圧器 (Efficient Multimodal Transformer) を提案する。
論文 参考訳(メタデータ) (2022-08-16T08:02:30Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。