論文の概要: The Interlingua Hypothesis: LLMs Translate via a Latent Task-agnostic Feature Space
- arxiv url: http://arxiv.org/abs/2609.00515v1
- Date: Tue, 01 Sep 2026 00:30:22 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-02 16:31:36.173459
- Title: The Interlingua Hypothesis: LLMs Translate via a Latent Task-agnostic Feature Space
- Title(参考訳): Interlingua仮説:LLMs Translate through a Latent Task-Agnostic Feature Space
- Authors: Jacob Brinton, Jannik Brinkmann, Mark Crovella, Aaron Mueller,
- Abstract要約: 大規模言語モデル(LLM)は、最近、強力な教師付きベースラインよりも機械翻訳性能が向上したことを示した。
近年の解釈可能性に触発されて,Interlingua仮説を提案する。
この仮説では、言語モデルは、原文を潜在特徴空間に読み込んで翻訳し、潜在特徴空間から読み出してターゲット文を生成する。
- 参考スコア(独自算出の注目度): 19.151700851322897
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Large language models (LLMs) have recently demonstrated improved machine translation performance over strong supervised baselines. This raises questions as to what mechanisms underlie how LLMs perform machine translation between languages. Motivated by recent interpretability findings--namely, that LLMs use massively multilingual latent feature representations to perform language modeling--we propose the interlingua hypothesis. The hypothesis holds that language models translate by reading a source sentence into a latent feature space, and generate a target sentence by reading from the latent feature space. We show three lines of evidence in support of this hypothesis: (1) variance in BLEU across language pairs is largely predictable from language-specific competences with no language pair-specific interaction terms; (2) many model components are causally influential in both monolingual tasks and translation tasks; and (3) fine-tuning on monolingual data recovers a large proportion of translation improvements relative to fine-tuning on aligned documents. Together, these provide convergent evidence in support of the interlingua hypothesis, and suggest new ways of understanding and improving how LLMs can be leveraged to perform translation tasks.
- Abstract(参考訳): 大規模言語モデル(LLM)は、最近、強力な教師付きベースラインよりも機械翻訳性能が向上したことを示した。
これにより、LLMが言語間の機械翻訳をどのように実行するのかという疑問が持ち上がる。
言語モデリングを行うためにLLMが多言語ラテント特徴表現を多言語で用いているという,近年の解釈可能性に関する知見に触発されて,我々はインターリングア仮説を提案する。
この仮説では、言語モデルは、原文を潜在特徴空間に読み込んで翻訳し、潜在特徴空間から読み出してターゲット文を生成する。
本仮説は,(1)言語対間のBLEUのばらつきが,言語対固有の相互作用項を持たない言語固有の能力から大きく予測可能であること,(2)単言語タスクと翻訳タスクの両方に因果的に影響を及ぼすモデル成分が多数存在すること,(3)単言語データの微調整は,整列文書の微調整と比較して翻訳改善のかなりの割合を回復することを示す。
これらとともに、インターリングア仮説を支持する収束した証拠を提供し、翻訳タスクの実行にLLMをどのように活用するかを理解し、改善する方法を提案する。
関連論文リスト
- When Meanings Meet: Investigating the Emergence and Quality of Shared Concept Spaces during Multilingual Language Model Training [57.230355403478995]
本研究では,EuroLLMの事前学習における言語に依存しない概念空間の開発について検討する。
共有概念空間は早期に出現し、洗練され続けていますが、それらとの整合性は言語に依存しています。
従来の作業とは対照的に、細かな手作業分析により、翻訳品質の顕著な向上は、行動の変化を反映していることが判明した。
論文 参考訳(メタデータ) (2026-01-30T11:23:01Z) - Language Surgery in Multilingual Large Language Models [39.66404344691661]
大規模言語モデル(LLM)はタスクや言語にまたがる顕著な一般化機能を示している。
本稿では, LLMにおける自然に出現する表現アライメント, 特に中層における表現アライメントについて検討する。
Inference-Time Language Control (ITLC) を提案する。
論文 参考訳(メタデータ) (2025-06-14T11:09:50Z) - The Rise and Down of Babel Tower: Investigating the Evolution Process of Multilingual Code Large Language Model [59.357993924917]
本研究では,大規模言語モデル(LLM)における事前学習過程における多言語機能の進化について検討する。
本稿では,LLMが新たな言語能力を習得する過程全体を記述したBabel Tower仮説を提案する。
本論文では,多言語コードLLMのための事前学習コーパスを最適化する新しい手法を提案する。
論文 参考訳(メタデータ) (2024-12-10T08:28:57Z) - LLM-based Translation Inference with Iterative Bilingual Understanding [52.46978502902928]
大規模言語モデル(LLM)の言語間機能に基づいた,新しい反復的バイリンガル理解翻訳法を提案する。
LLMの言語横断的能力により、ソース言語とターゲット言語を別々にコンテキスト理解することが可能になる。
提案したIBUTは、いくつかの強力な比較法より優れている。
論文 参考訳(メタデータ) (2024-10-16T13:21:46Z) - Crosslingual Capabilities and Knowledge Barriers in Multilingual Large Language Models [62.91524967852552]
大規模言語モデル(LLM)は、多言語コーパスの事前訓練のため、一般的に多言語である。
しかし、これらのモデルは言語間の対応する概念、すなわち言語を横断的に関連付けることができるだろうか?
本研究は,言語横断的タスクにおける最先端LLMの評価である。
論文 参考訳(メタデータ) (2024-06-23T15:15:17Z) - Empowering Cross-lingual Abilities of Instruction-tuned Large Language
Models by Translation-following demonstrations [0.8133739801185272]
We propose CrossAlpaca, a It-LLM with cross-lingual instruction-following and translation-following demonstrations。
我々のモデルは、6つの異なる言語でテストされ、単言語データで調整された It-LLM よりも優れています。
論文 参考訳(メタデータ) (2023-08-27T19:22:12Z) - Eliciting the Translation Ability of Large Language Models via Multilingual Finetuning with Translation Instructions [68.01449013641532]
大規模事前学習言語モデル(LLM)は多言語翻訳において強力な能力を示している。
本稿では,多言語事前学習言語モデルであるXGLM-7Bを微調整して,多言語翻訳を行う方法を提案する。
論文 参考訳(メタデータ) (2023-05-24T12:00:24Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。