論文の概要: Who Gets Named: Citation Type Predicts Individual Naming by Grounded Language Models, and a Roster Instrument Captures 0.5% of It
- arxiv url: http://arxiv.org/abs/2607.23893v2
- Date: Sat, 01 Aug 2026 21:12:53 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-04 21:56:46.32938
- Title: Who Gets Named: Citation Type Predicts Individual Naming by Grounded Language Models, and a Roster Instrument Captures 0.5% of It
- Title(参考訳): 名前の由来は: Citation typedicts individual Naming by Grounded Language Models, and a Roster Instruments Captures 0.5%
- Authors: Dmitrij Żatuchin,
- Abstract要約: 2026年7月24日に1つの2時間窓で2,400件の接地APIコールを発行した。
すべてのレスポンスは、個別のプロフェッショナルを指名するかどうかによってコード化されました。
モデルは25.8%の回答で個人を指名した。
- 参考スコア(独自算出の注目度): 0.0
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Prior work on AI brand visibility measures the firm: does a model recommend a company, and does that track its reputation. This study asks the question one level down, in categories where the buyer picks a person. It issued 2,400 grounded API calls in one two-hour window on 24 July 2026: 120 buyer-intent prompts, four models (GPT-5.6 Sol, Gemini 3.6 Flash, Perplexity Sonar Pro, Grok 4.5), five iterations each, four European markets and five query languages. Every response was coded for whether it named an individual professional, by a rule cascade that never consults a roster and that drops detections resolving to a same-named American city (precision 96.9%, recall 61.7%, so every rate below is a lower bound). All inference corrects for clustering within prompt: intraclass correlation 0.258, effective n 407 against a nominal 2,400. Models named an individual in 25.8% of responses. Category dominates: real estate 35.4% and car dealerships 32.9% against insurance 9.1% (chi-square 159.3, p = 5.8e-8 after correction). Models differ four-fold, from Grok 38.0% to Gemini 9.3%. Citation type predicts naming and citation volume does not: naming responses cite the individual's own site 2.6 points more often (95% CI +1.4 to +3.9) and category portals 4.3 points more often, and cite firm-owned pages at the same rate (44.1% against 45.5%). On nine matched translation pairs, English prompts named an individual in 36.7% of responses against 15.6% for the same question in the local language (OR 3.14, clustered p = 0.074, so the direction is clear and the design cannot close it). A 939-person roster built from public LinkedIn search matched 128 of 27,293 name-shaped mentions (0.47%), 26 of the 939 people were ever named, and the roster-derived rates of 0.0% to 25.4% measure that overlap. Roster-based measurement of individual AI visibility sees a small and unrepresentative slice of what models do.
- Abstract(参考訳): AIブランドの可視性に関する以前の研究は、会社を測る: モデルは会社を推薦し、その評判を追跡する。
この調査では、購入者が人を選別するカテゴリにおいて、質問を1レベル下げる。
2026年7月24日、120人の買い手によるプロンプト、4つのモデル(GPT-5.6 Sol、Gemini 3.6 Flash、Perplexity Sonar Pro、Grok 4.5)、各5つのイテレーション、4つのヨーロッパ市場、5つのクエリ言語。
すべての応答は、個々の専門家を指名するかどうかのコード化され、ロスターに相談せず、同じ名前のアメリカシティに解決する検出を落とすルールカスケードによって行われた(精度96.9%、リコール61.7%、下記のレートは下限である)。
クラス内相関0.258、n 407、名目2400に対して有効である。
モデルは25.8%の回答で個人を指名した。
カテゴリーは不動産35.4%、自動車ディーラー32.9%、保険9.1%(修正後159.3、p = 5.8e-8)である。
モデルはGrok 38.0%からGemini 9.3%と4倍に異なる。
命名応答は個人のサイト2.6ポイント(95% CI +1.4から+3.9)とカテゴリポータル4.3ポイント(45.5%に対して44.1%)を引用する。
9つの一致した翻訳対において、英語のプロンプトは、同じ質問に対する15.6%に対して36.7%の回答で個人を指名した(OR 3.14、クラスタ化されたp = 0.074、したがって方向は明確であり、設計はそれを閉じることができない)。
パブリックなLinkedIn検索で作られた939人のロスターは、27,293名の字型の言及(0.47%)のうち128名と一致し、939人の内26名が命名され、ロースター由来の比率は0.0%から25.4%と重複した。
Rosterベースの個々のAI可視性の測定は、モデルが何をしているかを小さく、表現できないスライスで見る。
関連論文リスト
- MemDelta: Controlled Baselines and Hidden Confounds in Agent Memory Evaluation [8.38765074765613]
エージェントメモリシステムは、RAGとフルコンテキストベースラインに対してますます評価される。
報告されたゲインは、しばしばメモリメソッドの変更と言語モデル、埋め込みモデル、または検索パイプラインの変更を混合する。
我々は,LongMemEval-S(500の質問,50以上のセッション,3つのモデルファミリ)において,一度に1つのコンポーネントを変更する制御された評価プロトコルであるMemDeltaを提案する。
論文 参考訳(メタデータ) (2026-06-29T07:51:22Z) - Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction [55.11480729304395]
PointArena 2026は77.2%の精度でベンチマークで2位である。
ap proachは3つの障害モードをターゲットにしている。第一に、エージェント駆動のシンセシスは大きなセマンティクスとアンカー相対的な候補プールを構築する。
次に、determinis tic steerable-dataパイプラインは、認証された10,000サンプルのメインセットと、マスク、テンプレート、パス検証を使用するリザーブサンプルを生成する。
論文 参考訳(メタデータ) (2026-06-29T06:39:03Z) - Knowledge Index of Noah's Ark [63.143852586221534]
KINAは,261分野にわたる899項目のベンチマークである。
ボーナス・オン・バートーナメントがFOSDを弱く支配していることを示す。
トップモデルであるGemini-3.1-Pro-Previewは53.17%、Claude-Opus-4.6は49.92%、GPT-5.4は48.55%に達した。
論文 参考訳(メタデータ) (2026-06-03T17:06:49Z) - FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks [0.0]
MathCheckのパラフレーズ品質検査では, 19群で4つの意味的不正確なパラフレーズが検出された。
GPT-4oは2位から4位へと降格し、クロード・ハイクとディープ・シークV3が上昇する。
論文 参考訳(メタデータ) (2026-05-27T18:59:18Z) - Evaluating Commercial AI Chatbots as News Intermediaries [85.32040752972836]
ベストシステムは、数時間前に報告されたイベントに関する質問に対して、90%以上の多重選択精度を達成する。
すべてのモデルはヒンディー語で最小の精度を達成する。
原因ではなく検索は エラーの70%以上を 引き起こします
論文 参考訳(メタデータ) (2026-05-21T17:42:07Z) - The Range Shrinks, the Threat Remains: Re-evaluating LLM Package Hallucinations on the 2026 Frontier-Model Cohort [51.56484100374058]
Spracklenらは、コード生成された大きな言語モデルは、PyPIやnpmに存在しないパッケージ名を幻覚させることを示した。
199,845対のPythonとJavaScriptプロンプトの幻覚率を測定し、PyPIとnpmマスターリストに対して検証した。
127個のパッケージ名(PyPIは109個,npmは18個)を5つの評価モデルで同一に作成する。
論文 参考訳(メタデータ) (2026-05-16T16:08:52Z) - Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation [0.0]
連鎖忠実性に関する最近の研究は、単一集合数について報告している。
本論文は、忠実性はモデルの客観的かつ測定可能な性質ではないことを示す。
論文 参考訳(メタデータ) (2026-03-20T17:48:43Z) - Distortions in Judged Spatial Relations in Large Language Models [45.875801135769585]
GPT-4は55%の精度で優れた性能を示し、GPT-3.5は47%、Llama-2は45%であった。
モデルは、ほとんどの場合において最も近い基数方向を同定し、その連想学習機構を反映した。
論文 参考訳(メタデータ) (2024-01-08T20:08:04Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。