論文の概要: Multilingual GSM-Symbolic: What determines capability transfer across languages?
- arxiv url: http://arxiv.org/abs/2610.03367v2
- Date: Mon, 05 Oct 2026 11:05:51 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-07 04:43:28.543443
- Title: Multilingual GSM-Symbolic: What determines capability transfer across languages?
- Title(参考訳): 多言語GSM-Symbolic: 言語間の能力伝達を決定するものは何か?
- Abstract要約: 我々は多言語GSM-Symbolicを用いて言語間能力伝達を評価する。
我々は,モデルのサイズ,言語資源レベル,推論,類型的距離など,能力の最大の決定要因を定量化する。
私たちの発見は、モデル開発者にとって重要な意味を持つ。
- 参考スコア(独自算出の注目度): 9.109101407022886
- License: http://creativecommons.org/licenses/by-sa/4.0/
- Abstract: We understand little about how capabilities acquired in one language carry over to another, or what governs this transfer: evaluations rely on incomparable, saturation-prone datasets and rarely examine its determinants jointly. Identifying what predicts transfer would let us avoid exhaustive evaluation across all language pairs and let developers target the factors that limit performance in low-resource languages. To evaluate cross-lingual capability transfer, we introduce Multilingual GSM-Symbolic, an extensible multilingual mathematical dataset covering 30,000 item-matched question-answer pairs and spanning 15 languages. It utilises symbolic templates to prevent overfitting and ensure generalisation by allowing generation of millions of high-quality variations from a single sample. Using Multilingual GSM-Symbolic, we quantify the largest determinants of capability as model size ($β= 1.77$), language resource level ($β= 0.77$), reasoning ($β= 0.67$) and typological distance ($β= -0.25$). This joint estimation allows these determinants to be expressed in terms of one another: a 32B model evaluated in Marathi performs like a 10B model in English. Our findings have important implications for model developers, showing that model size and reasoning narrow the performance gap between low- and high-resource languages ($β= -0.27$ and $β= -0.20$, respectively), while similar levers have little or no effect on typologically distant languages. Overall, our analysis framework explains 92% of between-language variation, but only 23% of the model-by-language variation, and predicts a model's performance on an unseen language within 6.0pp (r=.96). Incorporating measurements from just 10 templates in the target language reduces this to 4.19pp, enabling reasonable estimates of performance with little or no downstream dataset.
- Abstract(参考訳): 評価は互換性のない飽和確率データセットに依存しており、その決定要因を共同で調べることはめったにない。
トランスファーの予測を特定すれば、すべての言語ペアに対する徹底的な評価を回避でき、低リソース言語のパフォーマンスを制限する要因を開発者がターゲットできるようになります。
GSM-Symbolicは3万項目の問合せを対象とし、15言語にまたがる拡張可能な多言語数学的データセットである。
シンボリックテンプレートを使用して、オーバーフィッティングを防ぎ、単一のサンプルから数百万の高品質なバリエーションを生成することで、一般化を保証する。
Multilingual GSM-Symbolic を用いて、最大機能決定因子をモデルサイズ(β= 1.77$)、言語リソースレベル(β= 0.77$)、推論(β= 0.67$)、タイプ的距離(β= -0.25$)として定量化する。
マラソンで評価された32Bモデルは、英語で10Bモデルのように機能する。
モデルディベロッパには重要な意味があり、モデルサイズと推論が低リソース言語(β=-0.27$と$β=-0.20$)のパフォーマンスギャップを狭めていることが示されている。
全体として、我々の分析フレームワークは言語間の変動の92%を説明しているが、モデルごとの変化の23%しか説明せず、6.0pp (r=.96) 以内の未確認言語におけるモデルの性能を予測している。
ターゲット言語で10のテンプレートを組み込むことで、これを4.19ppに削減し、ダウンストリームデータセットをほとんどあるいは全く使用せずに、パフォーマンスの合理的な見積を可能にする。
関連論文リスト
- M-GATE: Multilingual Grammar, Accuracy in Translation, and Efficiency Benchmark for Large Language Models [0.0]
M-GATE(英語: M-GATE)は、30言語にまたがる言語能力のベンチマークである。
80以上の構成で50以上のモデルを評価しました。
論文 参考訳(メタデータ) (2026-08-04T15:16:05Z) - What if I ask in \textit{alia lingua}? Measuring Functional Similarity Across Languages [9.29956338690412]
我々の分析によると、モデルのサイズと能力が大きくなるにつれて、言語間でモデルの応答がますます一貫性を増す。
これらの結果は、多言語信頼性を評価する実用的なツールとしての$kappa_p$の価値だけでなく、より一貫性のある多言語システムの開発を導く可能性も浮き彫りにしている。
論文 参考訳(メタデータ) (2025-09-04T09:08:39Z) - The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants [80.4837840962273]
私たちは122の言語変種にまたがるデータセットであるBelebeleを紹介します。
このデータセットは、高、中、低リソース言語におけるテキストモデルの評価を可能にする。
論文 参考訳(メタデータ) (2023-08-31T17:43:08Z) - Memory-efficient NLLB-200: Language-specific Expert Pruning of a
Massively Multilingual Machine Translation Model [92.91310997807936]
NLLB-200は202言語をカバーする多言語ニューラルネットワークモデルである。
そこで本研究では,最大80%のエキスパートの除去を,それ以上の微調整を行なわずに行うことができるプルーニング法を提案する。
論文 参考訳(メタデータ) (2022-12-19T19:29:40Z) - Detecting Languages Unintelligible to Multilingual Models through Local
Structure Probes [15.870989191524094]
我々は、言語間モデルでよく理解されていない言語を検出するために、未理解のテキストのみを必要とする一般的なアプローチを開発する。
我々のアプローチは、もしモデルの理解が言語のテキストに対する摂動に無関心であるなら、その言語について限られた理解を持つ可能性が高いという仮説から導かれる。
論文 参考訳(メタデータ) (2022-11-09T16:45:16Z) - Few-shot Learning with Multilingual Language Models [66.49496434282564]
多様な言語群をカバーするバランスの取れたコーパス上で,多言語の自動回帰言語モデルを訓練する。
私たちの最大のモデルは、20以上の代表言語で数ショットの学習において、新しい最先端の技術を定めています。
本稿では,モデルがどこで成功し,失敗するかを詳細に分析し,特に言語間の文脈内学習を可能にすることを示す。
論文 参考訳(メタデータ) (2021-12-20T16:52:35Z) - Inducing Language-Agnostic Multilingual Representations [61.97381112847459]
言語間の表現は、世界中のほとんどの言語でNLP技術が利用可能になる可能性がある。
i) 対象言語のベクトル空間をピボットソース言語に再配置すること、(ii) 言語固有の手段と分散を取り除くこと、(ii) 副産物としての埋め込みの識別性を向上すること、(iii) 形態的制約や文の並べ替えを除去することによって言語間の入力類似性を高めること、の3つのアプローチを検討する。
論文 参考訳(メタデータ) (2020-08-20T17:58:56Z) - XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning [68.57658225995966]
XCOPA (Cross-lingual Choice of Plausible Alternatives) は11言語における因果コモンセンス推論のための多言語データセットである。
提案手法は,翻訳に基づく転送と比較して,現在の手法の性能が低下していることを明らかにする。
論文 参考訳(メタデータ) (2020-05-01T12:22:33Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。