論文の概要: Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders
- arxiv url: http://arxiv.org/abs/2608.23809v1
- Date: Mon, 24 Aug 2026 20:17:55 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-26 14:09:34.576677
- Title: Discovering Cross-Language Reasoning Invariance in LLMs with Geometry-Invariant Sparse Autoencoders
- Title(参考訳): 幾何学不変スパースオートエンコーダを用いたLLMにおけるクロスランゲージ推論不変性の検出
- Authors: Igor Bogdanov, Changcheng Huang,
- Abstract要約: 多言語言語モデルが類似した出力しか生成しない共有機能や言語固有の計算に依存しているかどうかを考察する。
このサンプルでは、言語間の機能共有がモデルに依存しており、異なるモデルで異なる深さに現れる。
- 参考スコア(独自算出の注目度): 0.7555681563742543
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Multilingual language models can solve the same mathematical problem in different languages, but it remains unclear whether they rely on shared features or on language-specific computations that only produce similar outputs. We study this question in five models from four families using the Multilingual Grade School Math (MGSM) dataset, with problems solved in English, German, French, Spanish, Russian, and Chinese, retaining problems with valid reasoning traces in all six languages and replaying those traces through the model to record representations at multiple layers. For each model, we first use Centered Kernel Alignment (CKA) to identify layers with cross-language alignment. At each selected layer, we train two sparse autoencoders (SAE): a baseline reconstruction-only model and a contrastive variant introduced in this work, the Geometry-Invariant SAE (GI-SAE). GI-SAE supplements the reconstruction loss with an Information Noise-Contrastive Estimation (InfoNCE) loss that trains the encoder to produce similar activations for traces of the same problem, regardless of language or token position. We then test whether the resulting shared features are functionally interchangeable by swapping their values between languages during the model's forward pass and measuring the resulting change in output, quantified by Kullback-Leibler (KL) divergence per feature. Although GI-SAE yields higher CKA and Jaccard similarity at nearly every layer, higher geometric similarity does not consistently imply greater functional interchangeability. We find that cross-language feature sharing is model- and architecture-dependent in this sample and appears at different depths in different models. GI-SAE primarily amplifies cross-language structure already present: the pattern is model-specific, with strengthening in Qwen, no functional benefit in Gemma, and mixed layer-dependent effects in Llama and Phi.
- Abstract(参考訳): 多言語言語モデルは、異なる言語で同じ数学的問題を解くことができるが、それらが共通の特徴や類似の出力しか生成しない言語固有の計算に依存しているかどうかは不明だ。
本研究は, 英語, ドイツ語, フランス語, スペイン語, ロシア語, 中国語で解決された問題と, 6言語すべてで有効な推論トレースの問題と, モデルを通してそれらのトレースを再生し, 複数の層での表現を記録することを目的とした, MGSM(Multilingual grade School Math)データセットを用いた4つのモデルの5つのモデルを用いて検討した。
それぞれのモデルに対して、まずCentered Kernel Alignment(CKA)を使用して、言語間のアライメントを持つレイヤを特定します。
選択された各層において、2つのスパースオートエンコーダ(SAE:sparse autoencoder)を訓練する:ベースライン再構成のみのモデルと、この研究で導入された対照的な変種であるGeometry-Invariant SAE(GI-SAE)である。
GI-SAEは、言語やトークンの位置に関わらず、同じ問題のトレースに対して同様のアクティベーションを生成するようにエンコーダを訓練するインフォメーションノイズコントラスト推定(InfoNCE)損失で再構築損失を補う。
次に、モデルの前方通過中に言語間でそれらの値を交換し、結果の出力変化を測定することで、結果の共有機能が機能的に交換可能であるかどうかを、Kullback-Leibler (KL) によって定量化する。
GI-SAEは、ほぼ全ての層において高いCKAとジャカードの類似性をもたらすが、高い幾何学的類似性は、常に機能的交換性を示すものではない。
このサンプルでは、言語間の機能共有がモデルに依存しており、異なるモデルで異なる深さに現れる。
GI-SAE は、Qwen の強化、Gemma の機能的メリット、Llama と Phi の混合層依存効果など、既に存在する言語間構造を増幅する。
関連論文リスト
- Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders [51.380449540006985]
大規模言語モデル(LLM)は多くの言語を処理できるが、どのようにして内部的にこの多様性を表現しているのかは不明だ。
言語固有のデコーディングと多言語表現を共有できるのでしょうか?
層間トランスコーダ(CLT)と属性グラフを用いて内部メカニズムを解析する。
論文 参考訳(メタデータ) (2025-11-13T22:51:06Z) - Semantic Convergence: Investigating Shared Representations Across Scaled LLMs [4.172347145536457]
大きな言語モデルは、サイズの違いにもかかわらず、世界全体を広く類似した解釈可能な特徴に彫り込み、クロスモデル解釈の基盤として普遍性を補強する。
予備実験では、単一トークンからマルチトークン部分空間への解析を拡張し、意味論的に類似した部分空間が言語モデルと同様に相互作用することを示す。
論文 参考訳(メタデータ) (2025-07-21T07:09:32Z) - The Unreasonable Effectiveness of Model Merging for Cross-Lingual Transfer in LLMs [45.08958917457921]
大規模言語モデル(LLM)は、ハイソース言語以外のタスクで依然として苦戦している。
本研究では,タスク固有のポストトレーニングデータが不足している低リソース言語への言語間移動について検討する。
論文 参考訳(メタデータ) (2025-05-23T20:28:31Z) - Examining Multilingual Embedding Models Cross-Lingually Through LLM-Generated Adversarial Examples [29.62231663945077]
本稿では,並列文のみを必要とする軽量な評価タスクである言語間セマンティック識別(D)と,対向的気晴らしを生成するLarge Language Model(LLM)を導入する。
CLSDは、意味的に誤解を招くが、語彙的に類似した代替品の上に、真の並列文をランク付けする埋め込みモデルの能力を測定する。
我々の実験では、検索タスクに微調整されたモデルは、英語をピボットすることの恩恵を受ける一方、bitextマイニングモデルは、直接言語間設定で最高のパフォーマンスを示す。
論文 参考訳(メタデータ) (2025-02-12T18:54:37Z) - A Variational Hierarchical Model for Neural Cross-Lingual Summarization [85.44969140204026]
言語間の要約(英: cross-lingual summarization)とは、ある言語の文書を別の言語の要約に変換することである。
CLSに関する既存の研究は主にパイプライン手法の利用やエンドツーエンドモデルの共同トレーニングに重点を置いている。
条件付き変分自動エンコーダに基づくCLSタスクの階層モデルを提案する。
論文 参考訳(メタデータ) (2022-03-08T02:46:11Z) - Examining Scaling and Transfer of Language Model Architectures for
Machine Translation [51.69212730675345]
言語モデル(LM)は単一のレイヤのスタックで処理し、エンコーダ・デコーダモデル(EncDec)は入力と出力の処理に別々のレイヤスタックを使用する。
機械翻訳において、EncDecは長年好まれてきたアプローチであるが、LMの性能についての研究はほとんどない。
論文 参考訳(メタデータ) (2022-02-01T16:20:15Z) - A Unified Strategy for Multilingual Grammatical Error Correction with
Pre-trained Cross-Lingual Language Model [100.67378875773495]
本稿では,多言語文法的誤り訂正のための汎用的かつ言語に依存しない戦略を提案する。
我々の手法は言語固有の操作を使わずに多様な並列GECデータを生成する。
NLPCC 2018 Task 2のデータセット(中国語)で最先端の結果を達成し、Falko-Merlin(ドイツ語)とRULEC-GEC(ロシア語)の競合性能を得る。
論文 参考訳(メタデータ) (2022-01-26T02:10:32Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。