論文の概要: Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories
- arxiv url: http://arxiv.org/abs/2609.01556v1
- Date: Tue, 01 Sep 2026 17:19:57 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-02 16:31:36.891727
- Title: Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories
- Title(参考訳): 検索されるがランク付けされない: 構造的検索における表面形式バイアス、数学からエージェント・トラジェクトリへ
- Authors: Nabira Rashid, Manolis Kellis,
- Abstract要約: 表面形状と意味を意図的に分離した埋め込み検索の評価を行った。
数学では、失敗は完全である: 最も重い偽りの階層における厳密なhit@1は、両方のプロダクション埋め込みに対して0.0%である。
表面の変化が付随する軌道では、同じモデルが超幾何学的確率で着陸する。
- 参考スコア(独自算出の注目度): 1.3092549225920291
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: We evaluate embedding retrieval where surface form and meaning are pulled apart on purpose: retrieving items that share underlying structure but not wording, in two unrelated domains under one protocol, competition mathematics (MathNet-Retrieve; 500 queries, 117,088-item corpus) and embodied-agent trajectories (ALFWorld-derived; 118 queries, 336 trajectories). In mathematics the failure is complete: strict Hit@1 at the heaviest disguise tier is 0.0% for both production embedders (bootstrap 95% CI [0.0, 0.0]) while the correct item sits in the top 10 nearly always, and in 95.2 to 99.8% of misses the winner is more lexically similar to the query than the correct answer. In trajectories, where surface variation is incidental, the same models land at or near hypergeometric chance when gold must involve a different object, and below chance for all three embedders once gold must differ in object and receptacle: retrieval anchors on literal tokens, not task structure. A lexical reranker control hurts in mathematics and helps in trajectories (closing 26 to 36% of the gap, CIs excluding zero); its sign reveals whether a benchmark's surface variation is adversarial or incidental. An LLM reranker recovers 5 to 63% of the gap in mathematics and 43 to 76% in trajectories; direction replicates across three judges (all 21 cells positive), but effect sizes, tier profiles, and the outlier judge change with domain (paired differences excluding zero everywhere). Mathematics gains concentrate on well-known competitions (+19.8 points, CI [+6.7, +33.2], one of six cells), so part of the recovery is memorization. In a paired downstream experiment (210 queries, graders at 96 to 99% agreement), oracle retrieval was indistinguishable from adversarially bad retrieval (McNemar p = 0.678); the solver's 69.5% zero-shot accuracy is largely a truncation proxy (97 to 100% on finished answers), leaving no headroom.
- Abstract(参考訳): 本研究では,2つの非関係領域(MathNet-Retrieve,500クエリ,117,088-item corpus)とエンボダイド・エージェント・トラジェクトリ(ALFWorld由来,118クエリ,336トラジェクトリ)を1つのプロトコルで検索する。
数学において失敗は完全である: 最も重い偽装階層における厳密なhit@1は、両方のプロダクション埋め込み(ブートストラップ95% CI [0.0, 0.0])に対して0.0%であり、正しい項目は、ほぼ常にトップ10にあり、95.2から99.8%のミスでは、勝者は、正しい答えよりもクエリによく似ている。
表面の変動が偶発的である軌跡では、同じモデルが、金が別の物体を巻き込まなければならず、また、一度金がオブジェクトとレセプタクルで異なる場合、3つの埋め込み器のチャンスより低いときに、超幾何学的確率で着陸する: タスク構造ではなく、リテラルトークンのアンカーを検索する。
レキシカルリランカ制御は数学の障害となり、軌跡(ギャップの26~36%、ゼロを除くCIを閉じる)に役立つ。
LLMリランカは数学のギャップの5~63%を回復し、トラジェクトリーでは43~76%を回復する。
数学はよく知られた競技(+19.8点、CI [+6.7, +33.2]、6つのセルのうちの1つ)に集中するので、回復の一部は記憶である。
対の下流実験(210クエリ、96~99%のグレーダー)では、オラクルの検索は逆の悪い検索とは区別できない(McNemar p = 0.678)。
関連論文リスト
- When Graph-JEPA Learns the Wrong Thing: Diagnosing and Repairing Category-Conditional Collapse [3.0255780640228984]
共同埋め込み予測アーキテクチャは線形探索と有効ランクによってほぼ普遍的に選択される。
本報告では,表現中に使用可能なインスタンス情報がゼロである場合に,両者が正常に読まれる事例を報告する。
Graph-JEPAは、サブグラフの残りの側面から1つのマスクされたアスペクトを予測し、線形プローブ精度0.871と有効ランク18-47を達成する。
再現性監査とターゲットゲートを備えたハーネスをリリースする。
論文 参考訳(メタデータ) (2026-08-20T19:22:49Z) - Efficient Visual Pointing for Embodied AI:Agent-Driven Data Synthesis, Cross-Block Attention, and Iterative Correction [55.11480729304395]
PointArena 2026は77.2%の精度でベンチマークで2位である。
ap proachは3つの障害モードをターゲットにしている。第一に、エージェント駆動のシンセシスは大きなセマンティクスとアンカー相対的な候補プールを構築する。
次に、determinis tic steerable-dataパイプラインは、認証された10,000サンプルのメインセットと、マスク、テンプレート、パス検証を使用するリザーブサンプルを生成する。
論文 参考訳(メタデータ) (2026-06-29T06:39:03Z) - Control-Plane Placement Shapes Forgetting: An Architectural Study of Agent Memory Across Thirteen System Configurations [1.2763567932588586]
ForgetEvalは1000ケースのテンプレート・スイートと385ケースの対向層(手作り+253 LLMのオラクル・バリデーション)
決定論的プリミティブは語彙的・時間的カテゴリーで十分だが、正準化に失敗する。
1000ケースのテンプレートスイートと385ケースの反対層であるForgetEvalを通じて、トレードオフを公開しています。
論文 参考訳(メタデータ) (2026-06-14T16:32:15Z) - Layer-Isolated Evaluation: Gating the Deterministic Scaffold of a Production LLM Agent with a No-LLM, Regression-Locked Test Harness [45.148328075418156]
デプロイされた注文エージェントは、階層の固定された分類に分解される。
純粋なスイートは、ロックされたスライス単位のベースラインに対して、すべての変更でCIで動作する。
制御されたレグレッションインジェクションにより、安全でない7つの層に一度に1つの層を分解し、検証する。
論文 参考訳(メタデータ) (2026-06-10T05:55:13Z) - Strained Coherence: A Pre-Failure Signal in Coding Agent Execution Trajectories [0.0]
LLMベースのコーディングエージェントは、時には自身の推論で問題を認識し、いずれにせよ前進する。
我々はClaude Sonnet 4.6のジャッジを構築し、完全なトラジェクトリとフラグがパターンの発生する場所にまたがる。
Qwen3.5-35B-A3Bのバックボーンを用いて44個の終端ベンチ2軌道上で評価を行った。
論文 参考訳(メタデータ) (2026-06-05T22:52:16Z) - Automated Proving of Shannon-Type Entropy Inequalities via Fine-Tuned Language Models and Guided Tree Search [50.16356451328644]
シャノン型エントロピーの不等式を証明することは情報理論の基本的な課題である。
我々は,原子実証のステップを微調整した小規模大規模言語モデルがこのプロセスを自動化することができるか検討する。
GPT-5.5は0ショットプロンプトで1.7%のサンプルを解き、Psitipは33.3%のサンプルを解いた。
論文 参考訳(メタデータ) (2026-06-04T05:43:12Z) - Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning [55.264863369127774]
現在の方法では、それぞれの正しいロールアウトを単一の報酬ビットに減らし、隠れた状態間で共有される幾何学的構造を無視している。
本稿では,RLトレーニングにおけるアンカートークンにおける正ロールアウトの最終層を,トレーニングと推論の両方においてゼロオーバーヘッドで整列する補助損失関数Hidden-Alignを提案する。
8つの数学的推論ベンチマークでは、Hidden-AlignはDAPOベースラインの平均パス@1をQwen3-1.7B, 4B, 14Bで3.8, 6.2, 5.4ポイント改善し、3つのスケールで一貫したパス@kゲインを得る。
論文 参考訳(メタデータ) (2026-06-02T06:51:15Z) - Evaluating Commercial AI Chatbots as News Intermediaries [85.32040752972836]
ベストシステムは、数時間前に報告されたイベントに関する質問に対して、90%以上の多重選択精度を達成する。
すべてのモデルはヒンディー語で最小の精度を達成する。
原因ではなく検索は エラーの70%以上を 引き起こします
論文 参考訳(メタデータ) (2026-05-21T17:42:07Z) - Less Is More: Cognitive Load and the Single-Prompt Ceiling in LLM Mathematical Reasoning [0.0]
SAIR Equational Theories Stage 1のコンペティションの文脈において,形式的数学的推論のためのプロンプトエンジニアリングについて検討する。
このタスクは、すべてのマグマに対して1つの方程式法則が別の法則を意味するかどうかを決定する必要がある。
5週間にわたって、40以上のプロンプトバリアントを設計、テスト、分析しました。
論文 参考訳(メタデータ) (2026-04-20T22:55:23Z) - LiveMathematicianBench: A Live Benchmark for Mathematician-Level Reasoning with Proof Sketches [61.30693283718321]
研究レベルの数学的推論のための動的多重選択ベンチマークであるLiveMathematicianBenchを提案する。
新たに発表された定理で評価を基礎づけることで、記憶されたパターンを超えた現実的なテストベッドを提供する。
このパイプラインは、高レベルな証明戦略を使用して、妥当だが無効な解選択を構築する。
論文 参考訳(メタデータ) (2026-04-02T08:22:17Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。