論文の概要: How Sparse Probability Maps Shape Mixture-of-Experts Routing
- arxiv url: http://arxiv.org/abs/2610.06677v1
- Date: Mon, 05 Oct 2026 16:47:20 GMT
- ステータス: 情報取得中
- システム内更新日: 2026-10-06 20:42:44.167614
- Title: How Sparse Probability Maps Shape Mixture-of-Experts Routing
- Title(参考訳): Sparse Probability map shape-of-experts Routing
- Abstract要約: Mixture-of-experts (MoE)ルータは通常、ルータのスコアにソフトマックスを適用して、トップKの専門家を維持する。
スパースマックス、アルファエントマックス、ノルムマックスのような空間誘導確率写像は、選択した専門家に正確な零点を適応的に割り当てることができる。
- 参考スコア(独自算出の注目度): 4.22444532909172
- License:
- Abstract: Mixture-of-experts (MoE) routers typically apply softmax to the router scores and keep the top-K experts, making every token use exactly K experts. Sparsity-inducing probability maps such as sparsemax, alpha-entmax and normmax can adaptively assign exact zeros to selected experts, and therefore appear to offer token-dependent expert participation, even when using the same top-K machinery. In this work, we study whether and how this sparsity survives training. We train matched 300M and 1B top-2 MoE language models with softmax, 1.5-entmax, sparsemax and 2-normmax, and find that the maps behave very differently once trained: at 1B, entmax discards 30% less probability mass than softmax while almost never dropping a selected expert, sparsemax retains the most mass, and normmax routes 21% of tokens to a single expert. These outcomes are not properties of the maps alone. Each map drops a selected expert only when the gap between the two largest scores reaches a fixed threshold, and the trained routers differ in the score distribution they learn: the entmax router learns scores with roughly half the spread of softmax's, which keeps its top-2 gaps below its threshold, while sparsemax and normmax, which share the same threshold, learn different gap distributions and hence different participation. Routers thus co-adapt their scores to the map, and a map's capacity to produce zeros does not by itself determine expert participation. While none of the sparse maps improves validation loss over softmax, they make the trained models far less sensitive to selecting more experts at inference: sparsemax trained with K=2 loses 0.02 nats when run with K=8, where softmax loses 0.58. Our results indicate that adaptive MoE routing has to be designed around the joint behavior of the probability map and the learned scores, rather than around the map alone.
- Abstract(参考訳): Mixture-of-experts (MoE)ルータは通常、ルータのスコアにソフトマックスを適用して、トップKの専門家を維持する。
スパースマックス(英語版)、アルファエントマックス(英語版)、ノルムマックス(英語版)のようなスパシティ誘導確率写像は、選択した専門家に正確なゼロを適応的に割り当てることができるため、同じトップK機械を使用してもトークン依存の専門家参加を提供するように見える。
本研究では,この空間がトレーニングを生き残るかどうかについて検討する。
私たちは300Mと1Bのトップ2 MoE言語モデルにソフトマックス、1.5-entmax、スパースマックス、および2-normmaxでマッチングし、1Bでは、エントマックスはソフトマックスよりも30%少ない確率の質量を捨てる一方、スパースマックスは最も多くの質量を保持し、ノルムマックスは1つの専門家にトークンの21%をルートする。
これらの結果は、写像のみの性質ではない。
各マップは、二つの大きなスコア間のギャップが一定の閾値に達した場合にのみ選択された専門家を降ろし、訓練されたルータは、彼らが学習するスコア分布が異なる。 entmaxルータは、その閾値以下でトップ2のギャップを保持するSoftmaxの約半分でスコアを学習し、同じ閾値を共有するスパースmaxとノルムマックスは、異なるギャップ分布を学習し、異なる参加をもたらす。
したがって、ルータはそのスコアをマップに適応させ、地図のゼロを生成する能力は、それ自体が専門家の参加を決定するものではない。
K=2でトレーニングされたスパースマックスは、K=8で実行すると0.02ナットを失い、ソフトマックスは0.58を失う。
この結果から, 適応型MoEルーティングは, 確率マップと学習スコアのジョイントな振舞いを中心に設計する必要があることが示唆された。
関連論文リスト
- MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference [17.718564947470906]
MaskCoFTは、クロスエントロピー損失だけでルーターとエキスパートを訓練する。
トークン単位のエキスパートフェッチを23.7%、ベースモデルに対して10.1%削減する。
9つのベンチマークの平均精度は、ベースモデルよりも0.92と0.53ポイント高い。
論文 参考訳(メタデータ) (2026-09-28T01:12:30Z) - Convergence Rates for Softmax Gating Mixture of Experts [78.3687645289918]
機械学習モデルの効率性とスケーラビリティを向上するための効果的なフレームワークとして、Mixture of Expert (MoE)が登場した。
MoEの成功の中心は、適応的なソフトマックスゲーティングメカニズムであり、各専門家の入力に対する関連性を決定する責任を負い、それぞれの重みを動的に専門家に割り当てる。
標準ソフトマックスゲーティングまたはその変種を備えたMoEの下で,パラメータ推定と専門家推定の収束解析を行い,密度とスパースゲーティングと階層ソフトマックスゲーティングを含む。
論文 参考訳(メタデータ) (2025-03-05T06:11:24Z) - MultiMax: Sparse and Multi-Modal Attention Learning [60.49318008131978]
SoftMaxは現代の機械学習アルゴリズムのユビキタスな成分である。
分散性はSoftMaxの変種族によって達成できるが、それらはしばしば代替損失関数を必要とし、多重モダリティを保たない。
入力入力範囲に応じて出力分布を適応的に変調するMultiMaxを提案する。
論文 参考訳(メタデータ) (2024-06-03T10:51:43Z) - Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts [78.3687645289918]
我々は,シグモイドゲーティング関数が,専門家推定の統計的タスクにおいて,ソフトマックスゲーティングよりも高いサンプル効率を享受できることを示した。
ReLU や GELU のようなよく使われる活性化型フィードフォワードネットワークとして定式化された専門家は,シグモイドゲーティングの下でより速い収束率を享受できる。
論文 参考訳(メタデータ) (2024-05-22T21:12:34Z) - r-softmax: Generalized Softmax with Controllable Sparsity Rate [11.39524236962986]
本稿では,ソフトマックスの修正であるr-softmaxを提案し,スパース確率分布を制御可能なスペーサ率で出力する。
我々は、r-softmaxが他のソフトマックス代替品よりも優れており、元のソフトマックスと高い競争力を持つ複数のマルチラベルデータセットを示す。
論文 参考訳(メタデータ) (2023-04-11T14:28:29Z) - Understanding Softmax Confidence and Uncertainty [95.71801498763216]
トレーニング分布から遠く離れたデータで予測する場合、ニューラルネットワークは不確実性を高めることができない、という指摘がしばしばある。
しかし、不確実性のプロキシとしてソフトマックスの信頼性を生かして、このためにのみテストするタスクにおいて、控えめな成功を達成します。
本稿では,この矛盾を解明し,ソフトマックスの信頼度と不確実性との相関を助長する2つの暗黙バイアスを同定する。
論文 参考訳(メタデータ) (2021-06-09T10:37:29Z) - Effectiveness of MPC-friendly Softmax Replacement [13.710300609457267]
我々は、ソフトマックス置換の2つの用途を分析し、ソフトマックスと比較する。
置換は1層ネットワークにおいて重要なスピードアップしか提供しないのに対して、常に精度を低下させ、時には著しく低下することがわかった。
論文 参考訳(メタデータ) (2020-11-23T04:14:32Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。