論文の概要: Permutation-Robust Decision Modeling with Candidate-Independent Block-Causal Attention
- arxiv url: http://arxiv.org/abs/2610.01601v1
- Date: Thu, 01 Oct 2026 12:47:58 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-03 01:19:24.133184
- Title: Permutation-Robust Decision Modeling with Candidate-Independent Block-Causal Attention
- Title(参考訳): カンジネート非依存なブロック因果性注意を用いた変分・ロバスト決定モデル
- Abstract要約: 決定モデルは、しばしば単一のシーケンスで符号化された可変サイズの候補アクションの集合をスコアする。
標準的な因果的クロスエンコーディングは表現力に富むが、候補のスコアを根底にある決定問題よりも順序に依存させることができる。
本稿では,共用コンテキスト内における因果計算を保存するため,候補非依存のブロック因果アテンションを導入する。
- 参考スコア(独自算出の注目度): 3.14496247732912
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Decision models often score a variable-sized set of candidate actions encoded in a single sequence. This setting is increasingly relevant for System 1 components inside generative systems, where candidates may be proposed or ordered differently across runs. Standard causal cross-encoding is expressive, but it can make a candidate's score depend on serialization order rather than on the underlying decision problem. We introduce candidate-independent block-causal attention, which preserves causal computation within the shared context and each candidate while blocking cross-candidate information flow and resetting candidate positions. We compare this architecture with standard causal attention and complementary invariant baselines across Gemma 3 1B, Qwen3 1.7B, and Qwen3 4B backbones. Candidate-independent attention consistently reduces permutation sensitivity while retaining competitive decision quality; ablations indicate that candidate isolation is the primary source of the effect, with position resetting completing the intended symmetry. A larger Qwen3-4B study further examines the behavior of the proposed architecture with substantially more training data. Code is available at the \href{https://github.com/guyAmit/ci-decision-models}{\textcolor{blue}{project repository}}, and the \href{https://huggingface.co/Guy-Amit/qwen3-4b-ci-decision-4096-poc}{\textcolor{blue}{Qwen3-4B model artifact}} is available on Hugging Face.
- Abstract(参考訳): 決定モデルは、しばしば単一のシーケンスで符号化された可変サイズの候補アクションの集合をスコアする。
この設定は、生成システム内のSystem 1コンポーネントにますます関係している。
標準的な因果的クロスエンコーディングは表現力に富むが、決定問題よりも直列化順序に頼らせることができる。
候補非依存のブロック因果的注意(ad candidate-independent block-causal attention)を導入し,各候補と共有コンテキスト内の因果的計算を保存するとともに,クロス候補情報フローをブロックし,候補位置をリセットする。
このアーキテクチャを、Gemma 3 1B、Qwen3 1.7B、Qwen3 4Bのバックボーンにまたがる標準的な因果的注意と相補的不変ベースラインと比較する。
候補非依存の注意は、競争的な決定品質を維持しながら、常に置換感度を低下させる。
より大規模なQwen3-4B研究は、より多くのトレーニングデータを用いて提案されたアーキテクチャの挙動をさらに調査する。
コードは、 \href{https://github.com/guyAmit/ci-decision-models}{\textcolor{blue}{project repository}}、 \href{https://huggingface.co/Guy-Amit/qwen3-4b-ci-decision-4096-poc}{\textcolor{blue}{Qwen3-4B model artifact}}で入手できる。
関連論文リスト
- QEncodeBench: Can Large Language Models Encode Classical Problems into Verified Quantum Oracles? [5.990509154718772]
グローバー探索、振幅増幅、および量子カウントは同じ再利用可能なサブルーチン(位相オラクル)に依存し、アルゴリズムの文献は与えられたように構成される。
QEncodeBenchは、古典的な制約問題を位相オラクルとしてエンコードした大きな言語モデルを処理します。
我々は、この能力を体系的に過大評価し、ベースステートテストをサンプリングし、系統的に過大評価する。
論文 参考訳(メタデータ) (2026-07-31T15:47:25Z) - Rushes: A Human Preference Dataset for Pluralistic Alignment [11.198313470367365]
Rushesは、対話的な物語環境における人間のエンゲージメントの嗜好を明らかにするためのデータセットとベンチマークである。
6試合で8,167人のユニークユーザーから44,226の意思決定イベントがある。
ユーザ選択は、一様ベースラインに対して低い選択エントロピーで定量化され、構造化された非ランダムパターンを示す。
論文 参考訳(メタデータ) (2026-07-22T22:32:34Z) - MBD: A Model-Based Debiasing Framework Across User, Content, and Model Dimensions [50.00784452900918]
この課題に対処する一般モデルベースデバイアス(MBD)フレームワークを提案する。
任意のコホートに対するエンゲージメント分布の文脈平均と分散を明示的に推定する。
この統合により、フレームワークはバイアス付き生信号からバイアスなしの表現に変換することができる。
論文 参考訳(メタデータ) (2026-03-15T15:07:01Z) - When Models Decide and When They Bind: A Two-Stage Computation for Multiple-Choice Question-Answering [7.622274098558385]
マルチチョイス質問応答(MCQA)は評価が容易だが、メタタスクを追加する。
本稿では,表現分析(PCA,線形プローブ)と因果介入を用いて,言語モデルがMCQAを内部的に実装する方法について検討する。
論文 参考訳(メタデータ) (2026-01-07T13:27:48Z) - Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions [103.20281438405111]
MCQA(Multiple-choice Question answering)は、高性能トランスフォーマー言語モデルのキーコンピテンスである。
我々は,正解を予測するための関連情報をエンコードするキー隠れ状態のローカライズに語彙予測とアクティベーションパッチ手法を用いる。
後続の層は語彙空間における予測応答記号の確率を増大させ、この確率の増加は、特異な役割を持つ注目ヘッドのスパースセットと関連していることを示す。
論文 参考訳(メタデータ) (2024-07-21T00:10:23Z) - Robust Outlier Rejection for 3D Registration with Variational Bayes [70.98659381852787]
我々は、ロバストアライメントのための新しい変分非局所ネットワークベース外乱除去フレームワークを開発した。
そこで本稿では, 投票に基づく不整合探索手法を提案し, 変換推定のための高品質な仮説的不整合をクラスタリングする。
論文 参考訳(メタデータ) (2023-04-04T03:48:56Z) - Quality-aware Part Models for Occluded Person Re-identification [77.24920810798505]
咬合は人体再識別(ReID)にとって大きな課題となる
既存のアプローチは一般的に、計算効率とReIDの精度の両面で最適であるように、目に見える身体の部品を推測するための外部ツールに依存している。
閉塞型ReIDのためのQPM(Quality-Aware Part Models)という新しい手法を提案する。
論文 参考訳(メタデータ) (2022-01-01T03:51:09Z) - Consensus-Guided Correspondence Denoising [67.35345850146393]
本稿では,地域間コンセンサス学習フレームワークと対応関係を異色化し,対応関係をロバストに識別する。
ローカル地域からグローバル地域への動的グラフから推定されるコンセンサススコアに基づいて,信頼度の高い候補を初期マッチングから蒸留する新しい「プルーニング」ブロックを導入した。
本手法は、堅牢なラインフィッティング、ワイドベースライン画像マッチング、画像ローカリゼーションベンチマークを顕著なマージンで上回る。
論文 参考訳(メタデータ) (2021-01-03T09:10:00Z) - Robust Question Answering Through Sub-part Alignment [53.94003466761305]
我々はアライメント問題として質問応答をモデル化する。
私たちは、SQuAD v1.1でモデルをトレーニングし、いくつかの逆および外ドメインデータセットでそれをテストします。
論文 参考訳(メタデータ) (2020-04-30T09:10:57Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。