論文の概要: Separating Syntax from Language: A Mechanistic Account of Translation in Multilingual LLMs
- arxiv url: http://arxiv.org/abs/2609.01356v1
- Date: Tue, 01 Sep 2026 14:58:38 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-02 16:31:36.786877
- Title: Separating Syntax from Language: A Mechanistic Account of Translation in Multilingual LLMs
- Title(参考訳): 言語から構文を分離する:多言語LLMにおける翻訳の力学的考察
- Authors: Mikhail Sonkin, Tanja Baeumel, Daniil Gurgurov, Josef van Genabith, Simon Ostermann,
- Abstract要約: 翻訳が以前想定されていたよりもモジュール化されていることを示す。
単語順の言語間差異を分離する多言語データセットを構築する。
我々は、構文変換に選択的に敏感な個々の注意ヘッドを同定する。
- 参考スコア(独自算出の注目度): 9.999248651696602
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Multilingual large language models (mLLMs) achieve strong performance in machine translation, yet our understanding of the mechanisms by which they transform representations from one language to another remains incomplete. Prior work suggests that translation decomposes into separable processes within an mLLM, where conceptual content is first represented independently, followed by a production into language-specific form. In this work, we show that translation is even more modular than previously assumed and that the output language production in translation processes is actually further separable into a syntax and a surface language process. We construct controlled multilingual datasets that isolate cross-linguistic differences in word-order and use causal interventions and probing to track how representations are transformed during translation. We find that models first construct target-side word-order before realizing the target language surface form. We identify individual attention heads that are selectively sensitive to syntactic transformations while remaining largely invariant to language identity. These results establish the commitment to a syntactic structure as an independent stage in translation, extending prior decompositions and showing how translation is implemented by functionally different components within mLLMs.
- Abstract(参考訳): 多言語大言語モデル(mLLM)は機械翻訳において高い性能を達成するが、表現をある言語から別の言語へ変換するメカニズムの理解はいまだ不完全である。
以前の研究では、翻訳はmLLM内で分離可能なプロセスに分解され、まず概念的内容が独立して表現され、次に言語固有の形式に生成される。
本研究は,翻訳が以前想定していたよりもモジュール化され,翻訳プロセスにおける出力言語生成が,構文と表面言語プロセスにさらに分離可能であることを示す。
我々は、単語順の言語間差異を分離する多言語データセットを構築し、因果介入を用いて、翻訳中に表現がどのように変換されるかを追跡する。
対象言語表層形式を実現する前に,まずターゲット側単語順を構築した。
構文変換に選択的に敏感な個別の注意ヘッドを識別するが、言語同一性はほとんど不変である。
これらの結果は、翻訳の独立した段階としての構文構造へのコミットメントを確立し、事前分解を拡張し、mLLM内の機能的に異なるコンポーネントによって翻訳がどのように実装されるかを示す。
関連論文リスト
- Finding the Translation Switch: Discovering and Exploiting the Task-Initiation Features in LLMs [69.28193153685893]
大きな言語モデル(LLM)は、タスク固有の微調整なしでも、しばしば強力な翻訳能力を示す。
このプロセスをデミスティフィケートするために、スパースオートエンコーダ(SAE)を活用し、タスク固有の特徴を特定するための新しいフレームワークを導入する。
我々の研究は、LLMの翻訳機構のコアコンポーネントをデコードするだけでなく、内部モデル機構を使用してより堅牢で効率的なモデルを作成するための青写真も提供しています。
論文 参考訳(メタデータ) (2026-01-16T06:29:07Z) - Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders [51.380449540006985]
大規模言語モデル(LLM)は多くの言語を処理できるが、どのようにして内部的にこの多様性を表現しているのかは不明だ。
言語固有のデコーディングと多言語表現を共有できるのでしょうか?
層間トランスコーダ(CLT)と属性グラフを用いて内部メカニズムを解析する。
論文 参考訳(メタデータ) (2025-11-13T22:51:06Z) - SASST: Leveraging Syntax-Aware Chunking and LLMs for Simultaneous Speech Translation [16.85064064077492]
本研究は,依存関係を解析することにより,入力ストリームを意味的に完全な単位に分割する文法に基づくチャンキング戦略を提案する。
SASST(Syntax-Aware Simultaneous Speech Translation)は,凍結したWhisperエンコーダとデコーダのみのLLMを統合したエンドツーエンドのフレームワークである。
論文 参考訳(メタデータ) (2025-08-11T09:13:35Z) - Evaluating Shortest Edit Script Methods for Contextual Lemmatization [6.0158981171030685]
現代の文脈補綴器は、単語の形式を補題に変換するために、しばしば自動的に誘導された短い編集スクリプト(SES)に依存している。
これまでの研究では,SESが最終補修性能にどのような影響を及ぼすかは調査されていない。
ケーシング操作と編集操作を別々に計算することは、全体として有益であるが、高機能な形態を持つ言語には、より明確に有用であることを示す。
論文 参考訳(メタデータ) (2024-03-25T17:28:24Z) - Decomposed Prompting for Machine Translation Between Related Languages
using Large Language Models [55.35106713257871]
DecoMTは、単語チャンク翻訳のシーケンスに翻訳プロセスを分解する、数発のプロンプトの新しいアプローチである。
DecoMTはBLOOMモデルよりも優れていることを示す。
論文 参考訳(メタデータ) (2023-05-22T14:52:47Z) - Examining Scaling and Transfer of Language Model Architectures for
Machine Translation [51.69212730675345]
言語モデル(LM)は単一のレイヤのスタックで処理し、エンコーダ・デコーダモデル(EncDec)は入力と出力の処理に別々のレイヤスタックを使用する。
機械翻訳において、EncDecは長年好まれてきたアプローチであるが、LMの性能についての研究はほとんどない。
論文 参考訳(メタデータ) (2022-02-01T16:20:15Z) - VECO: Variable and Flexible Cross-lingual Pre-training for Language
Understanding and Generation [77.82373082024934]
我々はTransformerエンコーダにクロスアテンションモジュールを挿入し、言語間の相互依存を明確に構築する。
独自の言語でコンテキストにのみ条件付けされたマスク付き単語の予測の退化を効果的に回避することができる。
提案した言語間モデルでは,XTREMEベンチマークのさまざまな言語間理解タスクに対して,最先端の新たな結果が提供される。
論文 参考訳(メタデータ) (2020-10-30T03:41:38Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。