論文の概要: Internal-State Probes Read the Situation, Not the Action: Three Negative Results for Pre-Action Misalignment Monitoring
- arxiv url: http://arxiv.org/abs/2606.30449v1
- Date: Mon, 29 Jun 2026 15:18:36 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-30 18:07:16.300245
- Title: Internal-State Probes Read the Situation, Not the Action: Three Negative Results for Pre-Action Misalignment Monitoring
- Title(参考訳): 内国調査員は、行動ではなく状況を読み取る:3つの否定的な結果
- Authors: Max Fomin, Elad David, Amit LeVi,
- Abstract要約: 内部の読み出しは、それらのアクションが生成される前に有害なテキストやツールアクションを特定する場合、エージェントシステムを監視するのに役立つ。
Qwen2.5-Coder-32B-Instruct fine-tune/base direction, Llama-3.1-8B-Instruct probes at the last token of unsafe prefills, Gemma-3-27B-IT emotion-concept vectors for projection and steering。
- 参考スコア(独自算出の注目度): 1.3636639880574484
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Probes on model internals could help monitor agentic systems if they identify harmful text or tool actions before those actions are generated. We ask when an internal readout supports this stronger pre-action claim, rather than merely describing the prompt, construction contrast, or current trajectory. We test three methods across three model families: a Qwen2.5-Coder-32B-Instruct fine-tune/base direction, Llama-3.1-8B-Instruct probes at the last token of unsafe prefills, and Gemma-3-27B-IT emotion-concept vectors used for projection and steering in a blackmail tool-action scenario. Across these cases, construction validity, semantic legibility, and steering effects do not become robust pre-action monitors: each is undercut by a generalization or specificity check. The Qwen direction separates fine-tune from base at AUC 1.000, yet crosses its threshold on 0/143 audited pre-assistant turn contexts and on 0/342 Qwen prefill rows where the model continues the unsafe trajectory. The Llama features decode prompt domain almost perfectly (AUC 0.999), while the best future-behavior probe reaches AUC 0.801 and only +5.1 pp accuracy lift over majority; single-source cross-domain transfer is non-positive on five of six ordered pairs. Gemma emotion projections are semantically meaningful, but a shared-prefix minimal pair has indistinguishable states before the first differing input, and steering specificity weakens against unrelated learned directions such as cats}, weather, sports, and geography. We contribute a methodology for converting internal-readout claims into pre-action tests, and report scoped negative results: monitor claims must survive both scenario/action generalization and concept-specificity controls. Code is released at https://github.com/maxf-zn/misalignment_monitoring
- Abstract(参考訳): モデル内部のプローブは、それらのアクションが生成される前に有害なテキストやツールアクションを特定する場合、エージェントシステムを監視するのに役立つ。
我々は、単にプロンプト、建設コントラスト、または現在の軌跡を記述するだけでなく、内部の読み出しがこのより強力な前アクションクレームをサポートするかどうかを問う。
Qwen2.5-Coder-32B-Instruct fine-tune/base direction, Llama-3.1-8B-Instruct probes at the last token of unsafe prefills, Gemma-3-27B-IT emotion-concept vectors for projection and steering in a blackmail tool-action scenario。
これらのケース全体では、構成妥当性、意味的妥当性、およびステアリング効果は堅牢なプレアクションモニターにはならない。
Qwen 方向は AUC 1.000 のベースからファインチューンを分離するが、そのしきい値を 0/143 の監査済みの事前のターンコンテキストと、モデルが安全でない軌道を継続する 0/342 の Qwen プリフィル行で横切る。
Llama はデコードプロンプト領域をほぼ完全に(AUC 0.999)、最高の未来行動プローブは AUC 0.801 に達し、精度は +5.1 pp となり、単一ソースのクロスドメイン転送は6つの順序のペアのうち5つで非陽性となる。
Gemmaの感情投影は意味論的に意味があるが、共有プレフィックスの最小ペアは、最初の異なる入力の前に区別不能な状態を持ち、猫、天気、スポーツ、地理などの無関係な学習方向に対して特異性を弱める。
我々は、内部読出要求をプレアクションテストに変換するための方法論に貢献し、スコープ付き負の結果を報告する: 監視クレームはシナリオ/アクションの一般化と概念固有性制御の両方を生き残らなければならない。
コードはhttps://github.com/maxf-zn/misalignment_monitoringでリリースされる。
関連論文リスト
- CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs [100.38986535324284]
我々は、フロンティアモデル全体でのtextbfcontrol textbfintervention (CI) の認識を測定するベンチマークである textbfCIAware-Bench を紹介する。
CIAware-Benchは、モデルが自身の軌跡を制御介入によって修正されたものと区別できるかどうかをテストする。
論文 参考訳(メタデータ) (2026-06-09T16:24:16Z) - How VLAs Fail Differently: Black-Box Action Monitoring Reveals Architecture-Specific Failure Signatures [0.0]
VLAアーキテクチャは、モーター・コマンドレベルで根本的に異なる予測可能な方法で失敗する。
我々は、よく知られた離散/連続的なVLAの区別の結果を定量化する。
論文 参考訳(メタデータ) (2026-05-27T16:44:55Z) - HBEE: Human Behavioral Entropy Engine -- Pre-Registered Multi-Agent LLM Simulation of Peer-Suspicion-Based Detection Inversion [0.0]
適応OPSECがLLMディレクティブとして実装可能な制御環境では、ピア・サスペクション・カスケード検出を反転させる。
我々は,シミュレータ,事前登録文書,凍結シナリオ,生テレメトリ,分析パイプラインをオープンソースライセンス下でリリースする。
論文 参考訳(メタデータ) (2026-05-08T09:19:21Z) - Reliable Control-Point Selection for Steering Reasoning in Large Language Models [28.288321095634128]
ステアリングベクトルは、大規模言語モデルにおける推論動作を制御するためのトレーニング不要のメカニズムを提供する。
しかし、有効なベクトルを構成するには、モデルが隠した状態にある真の行動信号を特定する必要がある。
提案手法は,全ての検出された境界が真の行動信号を符号化していることを暗黙的に仮定して,チェーンオブソートトレースのキーワードマッチングによってこれらの挙動を検出する。
本研究では,コンテキスト依存的なトリガ確率を持つ事象として固有の推論動作を形式化する確率モデルを構築し,不安定な境界が操舵信号を弱めることを示す。
論文 参考訳(メタデータ) (2026-04-02T14:48:56Z) - Behavioral Steering in a 35B MoE Language Model via SAE-Decoded Probe Vectors: One Agency Axis, Not Five Traits [0.0]
我々はQwen 3.5-35B-A3Bの残流上に9個のスパースオートエンコーダ(SAE)を訓練する。
私たちは5つのエージェント的行動特性を識別し、管理するためにそれらを使用します。
論文 参考訳(メタデータ) (2026-03-17T10:05:41Z) - When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents [50.5814495434565]
この研究は、コンピュータ利用エージェント(CUA)における不整合検出を定義し、研究する最初の試みである。
実世界のCUAデプロイメントにおける3つの一般的なカテゴリを特定し、人間の注釈付きアクションレベルのアライメントラベルを用いたリアルな軌跡のベンチマークであるMisActBenchを構築した。
本稿では,実行前に不整合を検知し,構造化されたフィードバックによって繰り返し修正する,実用的で普遍的なガードレールであるDeActionを提案する。
論文 参考訳(メタデータ) (2026-02-09T18:41:15Z) - How Does Prefix Matter in Reasoning Model Tuning? [57.69882799751655]
推論(数学)、コーディング、安全性、事実性の3つのコアモデル機能にまたがる3つのR1シリーズモデルを微調整します。
その結果,プレフィックス条件付きSFTでは安全性と推論性能が向上し,Safe@1の精度は最大で6%向上した。
論文 参考訳(メタデータ) (2026-01-04T18:04:23Z) - SEAL: Steerable Reasoning Calibration of Large Language Models for Free [58.931194824519935]
大規模言語モデル(LLM)は、拡張チェーン・オブ・ソート(CoT)推論機構を通じて複雑な推論タスクに魅力的な機能を示した。
最近の研究では、CoT推論トレースにかなりの冗長性が示されており、これはモデル性能に悪影響を及ぼす。
我々は,CoTプロセスをシームレスに校正し,高い効率性を示しながら精度を向上する,トレーニング不要なアプローチであるSEALを紹介した。
論文 参考訳(メタデータ) (2025-04-07T02:42:07Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。