論文の概要: Xeno-Interpretability: Investigating the Alien Minds of LLMs
- arxiv url: http://arxiv.org/abs/2609.20408v2
- Date: Mon, 21 Sep 2026 15:23:23 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-22 20:29:00.125927
- Title: Xeno-Interpretability: Investigating the Alien Minds of LLMs
- Title(参考訳): Xeno-Interpretability: LLMのエイリアン精神を探る
- Abstract要約: 本稿では、モデルが適切な人間の概念が存在しない区別を表現・活用できるかどうかを問う。
このような内部構造を xeno-representation と呼び、それらの研究を xeno-prepretability と呼ぶ。
Xeno-interpretabilityは、モデル内の人間の概念を見つけることから、モデル自体に固有の表現構造を発見し、特徴づけることへと、解釈可能性の目的をシフトさせる。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Large language models are usually interpreted through concepts that humans already possess: truthfulness, refusal, deception, personality, harmfulness, and related categories. This paper asks whether models may also represent and use distinctions for which no adequate human concept exists. We call such internal structures xeno-representations, and their study xeno-interpretability. We distinguish the human-interpretable semantic space from the xeno-semantic space: the region of model-native representations for which no adequate human conceptual counterpart is available. We show that the space of possible internal distinctions in an LLM is substantially larger than the space available through finite human descriptions. We then separate experimental identification from semantic interpretation: an internal representation may be reproducibly located, geometrically characterized, causally manipulated, and linked to downstream behaviour even when its semantic content cannot be adequately expressed in human terms. On this basis, we sketch an empirical programme to identify xeno-representations. We finally examine the implications for AI safety and multi-agent systems, where model-native representations may propagate and stabilize across interacting agents while remaining only partially visible through human-readable communication. Xeno-interpretability therefore shifts the aim of interpretability from finding human concepts inside models toward discovering and characterizing the representational structures that are native to the models themselves and might affect their behaviour in unpredictable ways.
- Abstract(参考訳): 大規模な言語モデルは通常、人間がすでに持っている概念によって解釈される:真実性、拒絶、偽り、個性、有害性、および関連するカテゴリ。
本稿では、モデルが適切な人間の概念が存在しない区別を表現・活用できるかどうかを問う。
このような内部構造を xeno-representation と呼び、それらの研究を xeno-prepretability と呼ぶ。
我々は、人間の解釈可能な意味空間と、モデルネイティブ表現の領域とを区別する。
LLMにおける内部の区別可能な空間は、有限人の記述によって得られる空間よりもかなり大きいことを示す。
内部表現は、人間の言葉で適切に表現できない場合でも、再現的に位置し、幾何学的に特徴付け、因果的に操作し、下流の振る舞いと関連付けられる。
そこで我々は,xeno-representationを識別するための経験的プログラムをスケッチする。
モデルネイティブ表現は、人間可読性通信によって部分的にのみ可視でありながら、相互作用するエージェント間で伝播し、安定化する可能性がある。
したがって、Xeno-interpretabilityは、モデルの中に人間の概念を見つけることから、モデル自体に固有の表現構造を発見し、特徴づけることへと、解釈可能性の目的をシフトさせ、予測不可能な方法でそれらの振る舞いに影響を与える可能性がある。
関連論文リスト
- Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding [96.81411333150213]
本稿では,最上位のMLLMが個別の意味空間をどのようにナビゲートするかを評価するためのベンチマークを紹介する。
モデルは基本的なシンボル認識に失敗することが多いが、複雑な推論タスクに成功している。
この作業は、より厳格で人間指向のインテリジェントなシステムを開発するためのロードマップを提供する。
論文 参考訳(メタデータ) (2026-03-19T04:08:20Z) - Revealing emergent human-like conceptual representations from language prediction [90.73285317321312]
大規模言語モデル(LLMs)は、人間らしい振る舞いを示すテキストの次のトーケン予測によってのみ訓練される。
これらのモデルでは、概念は人間のものと似ていますか?
LLMは、他の概念に関する文脈的手がかりに関連して、言語記述から柔軟に概念を導出できることがわかった。
論文 参考訳(メタデータ) (2025-01-21T23:54:17Z) - Perceptions of Linguistic Uncertainty by Language Models and Humans [26.69714008538173]
言語モデルが不確実性の言語表現を数値応答にどうマッピングするかを検討する。
10モデル中7モデルで不確実性表現を確率的応答に人間的な方法でマッピングできることが判明した。
この感度は、言語モデルは以前の知識に基づいてバイアスの影響を受けやすいことを示している。
論文 参考訳(メタデータ) (2024-07-22T17:26:12Z) - A Geometric Notion of Causal Probing [85.49839090913515]
線形部分空間仮説は、言語モデルの表現空間において、動詞数のような概念に関するすべての情報が線形部分空間に符号化されていることを述べる。
理想線型概念部分空間を特徴づける内在的基準のセットを与える。
2つの言語モデルにまたがる少なくとも1つの概念に対して、この概念のサブスペースは、生成された単語の概念値を精度良く操作することができる。
論文 参考訳(メタデータ) (2023-07-27T17:57:57Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。