論文の概要: Conflicting Supervision Moves Commitment, Not Capability: A 12.29σ arrangement effect that is exactly zero under a convention-agnostic score
- arxiv url: http://arxiv.org/abs/2610.00234v1
- Date: Wed, 23 Sep 2026 01:33:30 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-03 01:19:23.579991
- Title: Conflicting Supervision Moves Commitment, Not Capability: A 12.29σ arrangement effect that is exactly zero under a convention-agnostic score
- Title(参考訳): コンベンションスーパービジョン・モブズ・コミット, 能力の欠如: 12.29σアレンジメント効果は、慣例非依存スコアで正確にゼロである
- Abstract要約: 順序付けとスケジュールがそれぞれ異なる乗算因子として順序付け効果に入る境界を証明した。
パスが書いているのは、モデルがコミットする規約であり、正確なマッチベンチマークが見えない。
- 参考スコア(独自算出の注目度): 3.314947075424626
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: "Train a model on the same problems written under two incompatible conventions, both correct, and ask what the ordering of that data writes into the parameters. The learning-rate schedule is not a background condition for that question. It is the averaging operator, and it decides the answer. We prove a bound in which the arrangement and the schedule enter the ordering effect as separate multiplied factors: the arrangement only as a block period, the schedule only as how much weight the endpoint can place on any one moment of the run. A decaying schedule cannot put a large step size and an uncontracted remainder at the same moment; a constant one does exactly that at the last step. That decay moderates ordering effects has been reported in pretraining; the mechanism, the separation, and a controlled measurement of both halves are ours. Ten orderings of one corpus, one budget, everything but the path held fixed, run twice under families differing in lr_scheduler_type and nothing else: at a constant rate the interior spans 0.2221 in allocation, 11.63 contrast floors, monotone in how blocked the arrangement is. Under the single cosine every published arm uses, the same ten arms occupy two distinguishable states where their own resolution would allow about ten. "Order matters" and "order does not matter" are the two ends of one knob. What the path writes is which convention the model commits to, and no exact-match benchmark can see it. Across twelve arms acc_A+acc_B is constant to within 9.7% while the allocation share runs 0.04 to 0.87, so the 12.29-sigma arrangement switch this paper measures is exactly zero under a convention-agnostic metric. That conservation is quoted from the decayed family throughout, the constant-rate one being a noisier place to read it. Marking the convention in the prompt collapses the switch and reaches 87.5% of the union ceiling."
- Abstract(参考訳): 「整合性のない2つの慣行の下で記述された同じ問題を、そのデータの順序がパラメータに何に書き込まれるかというモデルにおいて、学習率のスケジュールは、その質問の背景条件ではない。平均演算子であり、答えを決定する。我々は、配列とスケジュールが別々の多重化要因として入る境界を証明している。配置はブロック期間のみであり、エンドポイントの重み付けは、実行中の任意の瞬間に設定できるだけである。崩壊スケジュールは、大きなステップサイズと未収縮の残りを同時に行うことはできない。定数は、最終段階において正確には、適度な順序付け効果が報告されている。メカニズム、分離、および制御された両者の順序付けは、すべて1回、固定されるが、1回、予算は1回、予算は1回、調整は1回、調整は1回、調整は1回、調整は1回、調整は1回、調整は1回、調整は1回、調整は1回、調整は1回、調整は1回、調整は1回、調整は1回、調整は1回、調整は1回、調整は1回、調整は1回、調整は1回、調整は1回、調整は1回、調整は0回、調整は1回、調整は0は0は1回、という条件は10回、という条件付き得る。」
関連論文リスト
- Measuring Collapse and Correction in Homogeneous-Panel LLM Debate [51.03982042769297]
マルチエージェント大言語モデル(LLM)の議論は、最終回答が改善するかどうかによってしばしば評価されるが、運動は必ずしも改善されない。
標準的な最終精度評価は、これらの反対のメカニズムを詳述する。
複数選択質問(MCQ)に関する同質な議論のための監査可能なプロトコルを導入し、各実行を、崩壊、修正、オンセット、署名された介入ユーティリティよりも、遷移台帳として記録する。
論文 参考訳(メタデータ) (2026-09-28T14:31:19Z) - Teacher-Anchored Selection of Post-Training Quantized Models under Domain Shift [46.60221265861393]
候補者と教師が固定され,対象ラベルが欠落または不足している家族に対する選択について検討した。
独立に訓練されたコンボリューション(英語版)またはビジョントランスフォーマー(英語版)の教師1人当たりの1人のうち1人、34人以上の候補者が、平均的な後悔を和らげる。
論文 参考訳(メタデータ) (2026-09-25T11:49:05Z) - Self-Cleaning and Captured Anyway: One Measured Primitive for Error in a Store an Agent Writes to Itself, and What a Falling Score Actually Measures [14.867869624585886]
「後から回収した店に結末を書き込むエージェントは、通常片道汚染として報告されるループを閉じる。」
「第87段の列は、全てアペンディックスWに、うち37段の行は退行、退行、失敗、自己是正、決定不能、または認められない限度であり、そうでないもの50である。」
論文 参考訳(メタデータ) (2026-09-06T13:54:02Z) - Shortcut Before Circuit: Document Statistics Time In-Context Conflict Resolution [0.0]
我々は、電流と希薄度が正確に一致した合成言語でトランスフォーマーを訓練する。
私たちはこれらを最小限の因果編集で分離し、真実、トークン数、答え位置を固定しながら1つのキューを反転させます。
論文 参考訳(メタデータ) (2026-08-25T12:09:19Z) - The Free-Recipe Limit: Every Recipe Effect Measures Which Premise of an Idealised Learner Broke [14.867869624585886]
本研究は,12モデル(0.5B-14B,3つの事前学習家族)で761個の微調整を施したレシピ検索壁を試作した。
固定体積の1つのコヒーレント領域内において、3つの古典的自由度は0.019階に対して0.010-0.021の平均となる。
公立のスコアカードは、事前登録された26のクレームを全て評価する: 18の支持、5の失敗、2の未試験、1の混合。
論文 参考訳(メタデータ) (2026-08-09T03:01:10Z) - Belief-Space Perception Routing under Coupled Sensor Faults and Compute Contention [0.0]
固定された時計を見て反応しなければならないロボットは、同時に2つの問題に遭遇する。
本稿では,センサフォア状態と計算競合状態の推定値を追跡する知覚ルータを提案する。
2つのストレッサーが共起している場合、結合されたポリシーは期限満了率を1.1から9.4ポイントに削減する。
論文 参考訳(メタデータ) (2026-07-31T22:17:12Z) - Phantom transitions in language model fine-tuning [0.0]
ほぼ同期の競合相手とコンテキスト上で言語モデルを微調整することは、しばしばサイレントに失敗する。
2つのファミリーにまたがる5つの変圧器アーキテクチャと5つのパラメータ範囲にまたがるこの構造について検討する。
位相遷移に類似した順序パラメータにおいて,鋭いカタパルト様ジャンプを観察する。
論文 参考訳(メタデータ) (2026-05-25T10:44:42Z) - Clustered Switchback Designs for Experimentation Under Spatio-temporal Interference [44.644520116360106]
我々は, 平均治療効果 (GATE) を推定し, 全単位を常に治療やコントロールに曝露した平均結果の差を推定した。
そこで我々は,単位をクラスタにグループ化し,時間ステップをブロックにグループ化する,クラスタ化されたスイッチバック設計を提案する。
良好なクラスタリングを許容するグラフに対して, トラッピングされたHorvitz-Thompson推定器が$tilde O(1/NT)$平均二乗誤差(MSE)を達成することを示す。
我々の結果は、citethu2022switchback、ugander2013graph、citetleung2022rateの結果を同時に一般化する。
論文 参考訳(メタデータ) (2023-12-25T01:00:58Z) - Problem Dependent View on Structured Thresholding Bandit Problems [73.70176003598449]
我々は、Thresholding Bandit problem (TBP)における問題依存体制について検討する。
学習者の目的は、シーケンシャルゲームの終わりに、所定のしきい値を超える手段を持つアームセットを出力することである。
コンケーブ設定と単調設定の両方で誤差の確率を上下に設定する。
論文 参考訳(メタデータ) (2021-06-18T15:01:01Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。