論文の概要: Attention Manifolds: Steering or Blocking Language Models by Editing Learned B-Spline Surfaces
- arxiv url: http://arxiv.org/abs/2610.00257v1
- Date: Thu, 24 Sep 2026 03:55:39 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-03 01:19:23.585977
- Title: Attention Manifolds: Steering or Blocking Language Models by Editing Learned B-Spline Surfaces
- Title(参考訳): 注意マニフォールド:学習したB-スプライン表面の編集によるステアリングまたはブロック言語モデル
- Abstract要約: この研究はテキスト・アテンション・多様体(texttextattention manifold)を導入し、クエリーキーの相互作用に基づいて各値次元を変調する2次元のB-スプライン曲面を学習する。
それぞれの表面はテンソル生成の立方体B-スプラインをゼロとし、事前訓練された挙動を保つ。
LLaMA 3.2-1B-Instruct と 3B-Instruct に適用すると、WikiText-2 の検証の難易度は 0.3% のパラメータオーバーヘッドで 2-2.5 ポイント削減される。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: In standard transformer attention, a source token sends the same value vector to every receiver. The query determines \emph{how much} to attend but not \emph{what} to extract. This work introduces \textbf{attention manifolds}: learned 2D B-spline surfaces $S_d(q_d, k_d)$ that modulate each value dimension based on the query-key interaction. Each surface is a tensor-product cubic B-spline initialized to zero, preserving pretrained behavior. Applied to LLaMA 3.2-1B-Instruct and 3B-Instruct, attention manifolds reduce WikiText-2 validation perplexity by 2--2.5 points with 0.3\% parameter overhead. Across 112 diverse prompts, surfaces change greedy-decoded output for 69\% (1B) to 83\% (3B) of cases, with the strongest effects on ambiguous and polysemous inputs (94--100\% change rate). The surfaces improve output quality: correcting factual errors (\emph{``the CAP theorem has three main components''} $\to$ \emph{``it is impossible to guarantee all three''}), increasing precision (\emph{``impossible to know certain properties''} $\to$ \emph{``impossible to know both position and momentum''}), and adding specificity (a generic quote $\to$ an attributed Saint Augustine citation, consistently at both scales). The learned surfaces are also mechanically editable: inverting a layer's coefficients changes greedy output for 9/10 prompts (KL~0.010), providing a geometric mechanism for model steering. Setting surface coefficients to $-1$ creates ``attention walls'' that block value flow through specific dimensions. In a preliminary experiment, a layer-wide wall redirects an explosive-device prompt from specific instructions to general educational content, suggesting a path toward safety-oriented manifold shaping.
- Abstract(参考訳): 標準変換器の注意では、ソーストークンは同じ値ベクトルをレシーバに送信する。
クエリは、参加すべき \emph{how much} を決定するが、抽出すべき \emph{what} ではない。
学習された 2D B-スプライン曲面 $S_d(q_d, k_d)$ は、クエリキーの相互作用に基づいて各値次元を変調する。
各表面は、ゼロに初期化されたテンソル生成の立方体B-スプラインであり、事前訓練された挙動を保つ。
LLaMA 3.2-1B-インストラクトと3B-インストラクトに適用すると、アテンション多様体はWikiText-2検証の難易度を0.3\%のパラメータオーバーヘッドで2-2.5ポイント削減する。
112の異なるプロンプトで、表面は69\% (1B) から83\% (3B) のグリーディ復号出力に変化し、その影響は曖昧で多義的な入力(94-100\%)に最も強い。
事実誤差の補正 ( CAP定理は3つの主成分''} $\to$ \emph{`It is impossible to guarantee all three'''})、精度の向上 (\emph{``impossible to know certain properties''} $\to$ \emph{`` Impossible to know both position and momentum'})、特異性の追加 (一般的な引用で$\to$ an attributeed Saint Augustine citation, consistent at both scales)。
層の係数を反転させると、9/10プロンプト(KL~0.010)のグリーディ出力が変化し、モデルステアリングの幾何学的なメカニズムが提供される。
表面係数を$1$に設定すると、特定の次元を流れる値をブロックする `<attention wall'' が生成される。
予備実験では、層幅の壁は、特定の指示から一般的な教育内容へ爆発装置のプロンプトをリダイレクトし、安全指向の多様体形成への道筋を示唆する。
関連論文リスト
- Harmfulness Propagation Dynamics: Layer-wise Trajectories of Adversarial Intent in Large Language Models [0.0]
textbfherald (textbfEncoding textbfRecognition via textbfActivation textbfLayer textbfDynamics)
textbfherald (textbfHarmful textbfEncoding textbfRecognition via textbfActivation textbfLayer textbfDynamics)
論文 参考訳(メタデータ) (2026-09-11T21:04:55Z) - Ockhamareto: Pareto-Gated Segment-Level Credit Assignment for Concise Unit-Test Generation with Reinforcement Learning [64.49364696628294]
我々は単発のGRPOフレームワークであるtextbfOckhamaretoを紹介した。
Ockhamaretoには2つの主要なコンポーネントがある: (i)emphPareto-gated Bonusは、(mutation, $-$#tests)スペースで非支配的なロールアウトのみを報酬する。
論文 参考訳(メタデータ) (2026-08-25T12:20:02Z) - Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors [51.56484100374058]
自動回帰モデルは、長時間のロールアウトでエラーを蓄積しますが、デプロイ時には、それを測定するための基本的な真実はありません。
我々は、方向フラグを介して動的システムを前方または後方にステップする単一の条件付き潜在拡散モデルを訓練する。
この双方向性は測定不要なテスト時間誤差信号を提供することを示す。
論文 参考訳(メタデータ) (2026-08-01T13:49:46Z) - SURGELLM: Rethinking Multi-Task Evaluation through Task-Aware Feature Gating with Class-Balanced Normalization [0.21485350418225238]
不均一なNLPタスクに展開される微調整エンコーダは、3つの複合的な問題に直面している。
textbfsurgellmは、専用の軽量モジュールでそれぞれに対処する統合トランスフォーマーフレームワークである。
論文 参考訳(メタデータ) (2026-06-23T07:47:21Z) - CaricHarmony: Contrastive Diffusion Paths for Identity-Preserving Caricature Synthesis [49.596677723190886]
スケッチベースの似顔絵合成は、基本的な失敗モードに悩まされる。
アイデンティティと形状の条件は拡散モデルに組み合わされ、地味な肖像画や認識不能な歪みに対して崩壊する。
並列な未汚染拡散経路を通じてこの汚染を明示的に解消する最初の訓練不要な手法であるCaricHarmonyを提案する。
論文 参考訳(メタデータ) (2026-06-11T22:57:59Z) - The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations [50.43168858368539]
大規模言語モデルは自信を持って時代遅れの回答を生成し、既存の方法では検出できない。
これは工学的な失敗ではなく構造的な失敗であり、時間的ドリフトは、幾何的に残留流の方向として、正確性と不確実性の両方に符号化される。
論文 参考訳(メタデータ) (2026-05-09T22:27:31Z) - Cascade Token Selection for Transformer Attention Acceleration [0.0]
カスケードメカニズムは、Layer $l$からLayer $l+1$に代表セットを継承し、$(T - r) times r$ cross-Gramで検証し、少数の追加と削除で更新する。
選択ステップのコストは、層ごとに$O(T2 d)$から$O(T r d)$に低下する。
論文 参考訳(メタデータ) (2026-05-04T19:49:13Z) - Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision Training [0.0]
変圧器における注意スコアは、低精度トレーニングにおけるオーバーフローリスクを最大で支配する2次形式である$S_ij = x_itop M x_j / sqrtd_h$である。
相互作用行列 $M = WQ WKtop$ が階数 $r ll d$ を持つとき、$max_i,j|S_ij|$ は $exp(-d22/) となる。
論文 参考訳(メタデータ) (2026-02-21T14:29:22Z) - Robust Layerwise Scaling Rules by Proper Weight Decay Tuning [50.11170157029911]
現代のスケール不変アーキテクチャでは、トレーニングは急速に劣化したグラデーション状態に入る。
我々は,AdamWに対して,幅をまたいだサブ層ゲインを保ったウェイトデカイスケーリングルールを導入する。
この結果は,パラメータが設定した定常スケールを明示的に制御することにより,ほぼ入出力体制を超えて$mu$Pを拡大する。
論文 参考訳(メタデータ) (2025-10-17T02:58:35Z) - Numerical Fragility in Transformers: A Layer-wise Theory for Explaining, Forecasting, and Mitigating Instability [0.0]
エラーがいつどこで発生するかを予測する一階のモジュールワイズ理論を提示する。
自己注意のために、3つの解釈可能な診断に分解する層間境界を導出する。
また、精度と幅を意識したLayerNormインジケータ$rho_rm LN$も導入する。
論文 参考訳(メタデータ) (2025-10-17T01:03:02Z) - On Understanding Attention-Based In-Context Learning for Categorical Data [49.40350941996942]
我々は,アテンションブロックで構成されるネットワークを開発し,各ブロックに自己注意層を付加し,その後にクロスアテンション層と関連するスキップ接続を付加する。
このモデルは、カテゴリー的観察を伴う文脈内推論のための多段階機能的GD推論を正確に行うことができる。
論文 参考訳(メタデータ) (2024-05-27T15:03:21Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。