論文の概要: What are They Thinking? Delineation, Probing and Tracking of Concepts in LLMs
- arxiv url: http://arxiv.org/abs/2605.28823v1
- Date: Tue, 07 Apr 2026 03:50:09 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-15 07:09:36.535573
- Title: What are They Thinking? Delineation, Probing and Tracking of Concepts in LLMs
- Title(参考訳): 彼らは何を考えているのか? LLMにおける概念の展開, 検証, 追跡
- Authors: Mohamed Abdelwahab, Michelle Yu Collins, Sihan Chen, Yi Cheng Zhao, Zafarullah Mahmood, Jiading Zhu, Soliman Ali, Jonathan Rose,
- Abstract要約: 広義の概念の集合の有無を検知するプローブの開発方法を示す。
これは4つの異なる概念と3つの異なるLLMで実現されている。
このプロセスがさらに多くの概念に拡張されると、新しいモデルを簡単に監視できるようになります。
- 参考スコア(独自算出の注目度): 5.802955093242262
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: As the influence of LLMs expands, it is imperative to gain insight into their decisions. One way to do that is to develop probes that detect the presence or absence of a broad set of concepts within the embeddings computed in an LLM - which is what we might say a model is "thinking" about. Such probes should be low-cost and easily applicable to any LLM, so that monitoring for many concepts is possible during normal operation. In this paper, we take the first steps towards developing the capability of creating many such probes by defining and executing examples of the key tasks needed: first, the careful delineation of a concept through the creation of a dataset with the concept both present and then absent. Then, the training and testing of a set of linear probes to detect the concept on any layer of an LLM, including an exploration of the complexity of the probe needed. Finally, we show that such probes can track concepts across larger contexts. This is done with four separate concepts and three different LLMs. When this process is scaled to many more concepts, it will create the ability to easily monitor new models.
- Abstract(参考訳): LLMの影響が拡大するにつれて、その決定に対する洞察を得ることが不可欠である。
その方法の1つは、LLMで計算された埋め込みの中に幅広い概念の存在や欠如を検出するプローブを開発することです。
このようなプローブは低コストで、どんなLLMにも容易に適用でき、通常の操作で多くの概念の監視が可能となる。
本稿では、まず、必要となる重要なタスクの例を定義し、実行することで、そのようなプローブを多数作成する能力を開発するための第一歩を踏み出す。
次に、LLMの任意の層における概念を検出するために一連の線形プローブの訓練と試験を行い、必要なプローブの複雑さの探索を含む。
最後に、そのようなプローブは、より大きなコンテキストをまたいだ概念を追跡できることを示す。
これは4つの異なる概念と3つの異なるLLMで実現されている。
このプロセスがもっと多くの概念に拡張されると、新しいモデルを簡単に監視できるようになります。
関連論文リスト
- Exploring Concept Depth: How Large Language Models Acquire Knowledge and Concept at Different Layers? [57.04803703952721]
大規模言語モデル(LLM)は、幅広いタスクで顕著なパフォーマンスを示している。
しかし、これらのモデルが様々な複雑さのタスクを符号化するメカニズムは、いまだに理解されていない。
概念深さ」の概念を導入し、より複雑な概念が一般的により深い層で得られることを示唆する。
論文 参考訳(メタデータ) (2024-04-10T14:56:40Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。