Fugu-MT 論文翻訳(概要): On Token's Dilemma: Dynamic MoE with Drift-Aware Token Assignment for Continual Learning of Large Vision Language Models

論文の概要: On Token's Dilemma: Dynamic MoE with Drift-Aware Token Assignment for Continual Learning of Large Vision Language Models

arxiv url: http://arxiv.org/abs/2603.27481v1
Date: Sun, 29 Mar 2026 02:30:55 GMT
ステータス: 翻訳完了
システム内更新日: 2026-03-31 23:18:44.979358
Title: On Token's Dilemma: Dynamic MoE with Drift-Aware Token Assignment for Continual Learning of Large Vision Language Models
Title（参考訳）: トケンのジレンマについて:大規模視覚言語モデルの継続的な学習のためのドリフト対応トケンアサインメントを用いた動的MoE
Authors: Chongyang Zhao, Mingsong Li, Haodong Lu, Dong Gong,
Abstract要約: ドリフト対応トークン代入でMoEを漸進的に拡張する動的MoEフレームワークを提案する。具体的には、トークンレベルのアサインガイダンスは、確立されたルーティングパターンを維持するために、新しい専門家から曖昧で古いトークンを分離する。我々のLLaVA-DyMoEは、ルーティングドリフトによって引き起こされる忘れを効果的に軽減し、平均的な最終精度で7%以上向上し、ベースラインと比較して忘れを12%減少させる。
参考スコア（独自算出の注目度）: 17.04431326257041
License: http://creativecommons.org/licenses/by/4.0/
Abstract: Multimodal Continual Instruction Tuning aims to continually enhance Large Vision Language Models (LVLMs) by learning from new data without forgetting previously acquired knowledge. Mixture of Experts (MoE) architectures naturally facilitate this by incrementally adding new experts and expanding routers while keeping the existing ones frozen. However, despite expert isolation, MoE-based continual learners still suffer from forgetting due to routing-drift: old-task tokens become mistakenly attracted to newly added experts, degrading performance on prior tasks. We analyze the failure mode at the token level and reveal the token's dilemma: ambiguous and old tokens in new-task data offer minimal learning benefit yet induce forgetting when routed to new experts, due to their ambiguous routing assignment during training. Motivated by this, we propose LLaVA-DyMoE, a dynamic MoE framework that incrementally expands the MoE with drift-aware token assignment. We characterize token types via their routing score distributions and apply targeted regularization. Specifically, a token-level assignment guidance steers ambiguous and old tokens away from new experts to preserve established routing patterns and alleviate routing-drift, while complementary routing score regularizations enforce expert-group separation and promote new-expert specialization. Extensive experiments demonstrate that our LLaVA-DyMoE effectively mitigates routing-drift-induced forgetting, achieving over a 7% gain in mean final accuracy and a 12% reduction in forgetting compared to baselines. The project page is https://zhaoc5.github.io/DyMoE.
Abstract（参考訳）: マルチモーダル・インストラクション・チューニングは、以前取得した知識を忘れずに新しいデータから学習することで、LVLM(Large Vision Language Models)を継続的に強化することを目的としている。 Mixture of Experts (MoE)アーキテクチャは、新たなエキスパートを段階的に追加し、ルータを拡大し、既存のアーキテクチャを凍結し続けることで、これを自然に促進します。しかし、専門家の隔離にもかかわらず、MoEベースの継続学習者は、ルーティング・ドリフトによる忘れがちである: 古いタスクトークンは、新しく追加された専門家に誤って惹かれ、以前のタスクのパフォーマンスが低下する。我々はトークンレベルでの障害モードを分析し、トークンのジレンマを明らかにする。新しいタスクデータの曖昧さと古いトークンは、トレーニング中のあいまいなルーティング割り当てのために、新しいエキスパートにルーティングされたときの忘れを誘発する、最小限の学習利益を提供する。そこで我々はLLaVA-DyMoEを提案する。LLaVA-DyMoEは動的MoEフレームワークで、ドリフト対応トークン代入でMoEを漸進的に拡張する。ルーティングスコア分布によってトークンの型を特徴付け、ターゲット正則化を適用する。具体的には、トークンレベルの割当てガイダンスは、確立されたルーティングパターンを維持し、ルーティング・ドリフトを軽減するために、新しい専門家から不明瞭で古いトークンを取り除き、補完的なルーティングスコアの正規化はエキスパートグループ分離を強制し、新しい専門家の専門化を促進する。我々のLLaVA-DyMoEは、ルーティングドリフトにより引き起こされる忘れを効果的に軽減し、平均的な最終精度で7%以上向上し、ベースラインと比較して忘れを12%削減することを示した。プロジェクトページはhttps://zhaoc5.github.io/DyMoE。

論文の概要: On Token's Dilemma: Dynamic MoE with Drift-Aware Token Assignment for Continual Learning of Large Vision Language Models

関連論文リスト