論文の概要: The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models
- arxiv url: http://arxiv.org/abs/2608.24662v2
- Date: Wed, 26 Aug 2026 16:10:10 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-27 14:15:15.226436
- Title: The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models
- Title(参考訳): Invisible Editorial Layer:Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models
- Authors: Augusto Camargo,
- Abstract要約: 生成言語モデルの評価は、しばしば、政治的スタンス、ブランドの傾き、規範的フレーミングといった観察可能な行動特性を、モデルの重みの表示、訓練後のアライメント、またはプロンプトとして解釈する。
モデルパラメータの変更を必要とせずに、組織的、イデオロギー的、商業的なフレームに対して生成されたテキストを体系的に操る。
生成システムのガバナンスは、モデルを監査することと、最終的に話すデプロイされたシステムの監査とを、ますます区別する必要がある、と我々は主張する。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Evaluations of generative language models frequently interpret observable behavioral traits, such as political stance, brand inclination, and normative framing, as manifestations of model weights, post-training alignment, or prompting. This interpretation risks conflating a foundation model with the multi-layered production system through which its outputs are ultimately served. Modern inference stacks support runtime interventions capable of modifying generation while model parameters remain frozen. We examine inference-time framing bias: systematic runtime steering of generated text toward institutional, ideological, or commercial frames without requiring changes to the underlying model parameters. We formalize the Inference Attribution Problem and establish an observational non-identifiability result showing that, under black-box observation alone, behaviorally equivalent deployed systems may arise from structurally distinct combinations of model parameters and inference policies. Consequently, observed behavioral bias does not uniquely identify the architectural layer responsible for it. We further characterize Probability Placement as a deployment pattern in which undisclosed commercial influence is embedded within an ostensibly organic assistant response through systematic probability-mass reallocation, distinguishing it from explicit token-auction mechanisms for generative advertising. Finally, we discuss implications for behavioral auditing, inference provenance, confidential computing, cryptographic attestation, the EU AI Act, the Digital Services Act, and advertising-disclosure principles. We argue that governance of generative systems must increasingly distinguish between auditing a model and auditing the deployed system that ultimately speaks.
- Abstract(参考訳): 生成言語モデルの評価は、しばしば、政治的スタンス、ブランドの傾き、規範的フレーミングといった観察可能な行動特性を、モデルの重みの表示、訓練後のアライメント、またはプロンプトとして解釈する。
この解釈は、ファンデーションモデルと、その出力が最終的に提供される多層生産システムとを混同するリスクを負う。
現代的な推論スタックは、モデルパラメータが凍結されている間、生成を変更可能なランタイム介入をサポートする。
モデルパラメータの変更を必要とせず、制度的、イデオロギー的、商業的フレームに対して生成されたテキストの体系的な実行時ステアリングについて検討する。
我々は、推論属性問題を定式化し、ブラックボックス観察単独で、モデルパラメータと推論ポリシーの構造的に異なる組み合わせから行動等価なデプロイシステムが生じることを示す観察的非識別性結果を確立する。
したがって、観察された行動バイアスは、その原因となるアーキテクチャ層をユニークに識別しない。
我々はさらに、確率配置を、確率質量の体系的な再配置を通じて、有機的アシスタント応答の中に、未公表の商業的影響が組み込まれている配置パターンとして特徴付け、生成広告のための明示的なトークン誘引メカニズムと区別する。
最後に、行動監査、推論証明、機密計算、暗号証明、EU AI法、デジタルサービス法、広告開示原則などについて論じる。
生成システムのガバナンスは、モデルを監査することと、最終的に話すデプロイされたシステムの監査とを、ますます区別する必要がある、と我々は主張する。
関連論文リスト
- Automatically Finding and Validating Unexpected Side-Effects of Interventions on Language Models [0.9099663022952497]
本研究では,大規模言語モデルに対する介入の行動的影響を評価するための,自動化されたコントラスト評価パイプラインを提案する。
我々は, 既知の行動変化を注入することにより, 合成条件下でのアプローチを評価し, パイプラインがそれらを確実に回復することを示す。
全体として、パイプラインは、介入によって引き起こされるモデル行動の変化のホック後の監査のための統計的に根拠付き、解釈可能なツールを提供する。
論文 参考訳(メタデータ) (2026-05-06T16:27:23Z) - Active Inference: A method for Phenotyping Agency in AI systems? [0.11904398364437437]
3つの基準を基準として、原則検査に開放された最小限の概念について論じる。
後続の信念、事前の嗜好、そして期待される自由エネルギーの最小化は、共同でエージェント・アクション・チェーンを構成する。
論文 参考訳(メタデータ) (2026-04-25T12:41:53Z) - Subject-Event Ontology Without Global Time: Foundations and Execution Semantics [51.56484100374058]
形式化は9つの公理(A1-A9)を含み、実行可能性の正しさを保証する:履歴の単調性(I1)、因果性の非巡回性(I2)、トレーサビリティ(I3)である。
フォーマル化は、分散システム、マイクロサービスアーキテクチャ、DLTプラットフォーム、およびマルチパースペクティビティシナリオ(異なる主題から事実を分解する)に適用できる。
モデルに基づくアプローチ(A9): スキーマによるイベント検証、アクター認可、グローバル時間なしで因果連鎖の自動構築(W3)。
論文 参考訳(メタデータ) (2025-10-20T19:26:44Z) - Cognition-of-Thought Elicits Social-Aligned Reasoning in Large Language Models [17.381122321801556]
大きな言語モデル(LLM)は複雑な推論において優れているが、それでも有害な振る舞いを示すことができる。
本稿では,認知的自己監視ループにLCMを組み込んだ新しい復号時間フレームワークCooTを紹介する。
論文 参考訳(メタデータ) (2025-09-27T18:16:57Z) - On the Fairness, Diversity and Reliability of Text-to-Image Generative Models [68.62012304574012]
マルチモーダル生成モデルは 信頼性 公正性 誤用の可能性について 批判的な議論を巻き起こしました
埋め込み空間におけるグローバルおよびローカルな摂動に対する応答を解析し、モデルの信頼性を評価するための評価フレームワークを提案する。
提案手法は, 信頼できない, バイアス注入されたモデルを検出し, 組込みバイアスの証明をトレースするための基礎となる。
論文 参考訳(メタデータ) (2024-11-21T09:46:55Z) - Interpretable Imitation Learning with Dynamic Causal Relations [65.18456572421702]
得られた知識を有向非巡回因果グラフの形で公開することを提案する。
また、この因果発見プロセスを状態依存的に設計し、潜在因果グラフのダイナミクスをモデル化する。
提案するフレームワークは,動的因果探索モジュール,因果符号化モジュール,予測モジュールの3つの部分から構成され,エンドツーエンドで訓練される。
論文 参考訳(メタデータ) (2023-09-30T20:59:42Z) - Answering Causal Queries at Layer 3 with DiscoSCMs-Embracing
Heterogeneity [0.0]
本稿では, 分散一貫性構造因果モデル (DiscoSCM) フレームワークを, 反事実推論の先駆的アプローチとして提唱する。
論文 参考訳(メタデータ) (2023-09-17T17:01:05Z) - Differentially Private Counterfactuals via Functional Mechanism [47.606474009932825]
本稿では,デプロイされたモデルや説明セットに触れることなく,差分的プライベート・カウンティファクト(DPC)を生成する新しいフレームワークを提案する。
特に、ノイズの多いクラスプロトタイプを構築するための機能機構を備えたオートエンコーダを訓練し、次に潜伏プロトタイプからDPCを導出する。
論文 参考訳(メタデータ) (2022-08-04T20:31:22Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。