Fugu-MT 論文翻訳(概要): NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles

論文の概要: NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles

arxiv url: http://arxiv.org/abs/2605.01847v1
Date: Sun, 03 May 2026 12:30:58 GMT
ステータス: 翻訳完了
システム内更新日: 2026-05-05 20:33:49.962959
Title: NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
Title（参考訳）: NeuroState-Bench: LLMエージェントプロファイルにおけるコミットインテリジェンスのためのヒューマンキャリブレーションベンチマーク
Authors: Jia Xiao,
Abstract要約: NeuroState-Benchは、ベンチマーク定義のサイドクエリープローブを通じてコミットメントの整合性を運用する、人間の校正ベンチマークである。主な32点評価は、固定された16点のローカルサブセットと、同一のベンチマークパイプラインで評価された16点のホストされた大型モデルサブセットを含む。経験的に、タスクの成功とコミットメントの整合性は、この拡張されたグリッドに分散します。
参考スコア（独自算出の注目度）: 0.4512372501420207
License: http://creativecommons.org/licenses/by/4.0/
Abstract: Outcome-only evaluation under-specifies whether an evaluated agent profile preserves the commitments required to solve a multi-turn task coherently. NeuroState-Bench is a human-calibrated benchmark that operationalizes commitment integrity through benchmark-defined side-query probes rather than inferred hidden activations. The released inventory contains 144 deterministic tasks and 306 benchmark-defined side-query probes spanning eight cognitively motivated failure families, paired clean and distractor variants, and three difficulty bands. The main 32-profile evaluation contains a fixed 16-profile local subset and a matched 16-profile hosted large-model subset evaluated through the same benchmark pipeline. Human calibration uses the final merged reporting scope: 104 sampled task units, 216 raw annotations, and 108 adjudicated task rows, with weighted kappa = 0.977 and ICC(2,1) = 0.977. Empirically, task success and commitment integrity diverge across this expanded grid: the success leader is not the integrity leader, 31 of 32 profiles change rank when integrity replaces task success, and integrity rankings are more stable under distractor perturbation. The primary confidence-free score HCCIS-CORE reaches 0.8469 AUC and 0.6992 PR-AUC for post-probe diagnostic discrimination of terminal task failure; the legacy full heuristic variant HCCIS-FULL reaches 0.7997 AUC and 0.6410 PR-AUC. Probe accuracy and state drift achieve slightly higher ROC-AUC, 0.8587, and better Brier/ECE, while HCCIS-CORE has substantially higher point-estimate PR-AUC and remains more closely tied to the benchmark's intended construct. The exploratory neural-augmented variant HCCIS+N is weaker overall, and a randomized subspace control approaches chance. NeuroState-Bench therefore contributes a calibrated evaluation axis for exposing commitment failures over a broader model grid than the original local-only subset.
Abstract（参考訳）: 結果のみの評価は、評価されたエージェントプロファイルがマルチターンタスクのコヒーレントな解決に必要なコミットメントを保存するかどうかを規定する。 NeuroState-Benchは、隠れたアクティベーションを推測するのではなく、ベンチマーク定義のサイドクエリープローブを通じてコミットメントの整合性を運用する人間校正ベンチマークである。リリースされたインベントリには、144の決定論的タスクと306のベンチマークで定義されたサイドクエリープローブが含まれており、8つの認知的なモチベーションを持つ障害ファミリー、ペアのクリーンとイントラクタのバリエーション、そして3つの困難バンドで構成されている。主な32点評価は、固定された16点のローカルサブセットと、同一のベンチマークパイプラインで評価された16点のホストされた大型モデルサブセットを含む。 104のサンプル化されたタスクユニット、216の生のアノテーション、108の調整されたタスク行で、加重されたkappa = 0.977とICC(2,1) = 0.977である。経験的に、タスクの成功とコミットメントの整合性は、この拡張されたグリッドに分散している。成功リーダは、整合性リーダーではない。 HCCIS-CORE は0.8469 AUC と 0.6992 PR-AUC に到達し、終末タスク障害の診断後診断が可能となり、旧来のフルヒューリスティック変種 HCCIS-FULL は0.7997 AUC と 0.6410 PR-AUC に到達した。精度と状態のドリフトは、ROC-AUC、0.8587、より優れたブライア/ECEを達成する一方、CCIS-COREは、かなり高い点推定PR-AUCを持ち、ベンチマークの意図した構成と密接に結びついている。探索的神経増強型HCCIS+Nは全体として弱く、ランダム化されたサブスペース制御がチャンスに近づいた。それゆえ、NeuroState-Benchは、元のローカルのみのサブセットよりも広いモデルグリッド上でのコミットメントの失敗を露呈するための、キャリブレーションされた評価軸に寄与する。

論文の概要: NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles

関連論文リスト