論文の概要: Task Structure Reverses Layerwise State Encoding in Sequence Models
- arxiv url: http://arxiv.org/abs/2606.00926v1
- Date: Sat, 30 May 2026 23:26:29 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-02 21:34:28.981002
- Title: Task Structure Reverses Layerwise State Encoding in Sequence Models
- Title(参考訳): Task Structure Reverses Layerwise State Encoding in Sequence Models
- Authors: Yuhang Jiang,
- Abstract要約: タスクが変更されたとき、同じアーキテクチャがこのプロファイルを反転させるのが分かります。
微調整されたMamba-130MとPythia-160Mでは、ParityはMambaとリカレントベースラインに徐々に集中し、Transformerによって徐々に建設された。
同じフリップが微調整されたMamba-130MとPythia-160Mに現れ、Pythia Dyckのボトルネックは410Mで持続する。
- 参考スコア(独自算出の注目度): 8.403971471573607
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Mechanistic studies of sequence models often treat layerwise state encodings as architectural traits: recurrent models concentrate readable state, attention-based models distribute it. We find that the same architecture reverses this profile when the task changes. Across Transformers, Mamba, Mamba-2, LSTMs, and GRUs, Parity is concentrated late in Mamba and the recurrent baselines and built gradually by Transformer; on bounded-depth Dyck-k the pattern flips. The same flip appears in fine-tuned Mamba-130M and Pythia-160M, and the Pythia Dyck bottleneck persists at 410M. Two explanations are conflated in the literature: algebraic structure (commutativity) versus computational structure (prefix update vs. stack). To separate them we add a third task: non-commutative S_3 permutation composition. S_3 groups with Parity, not Dyck, on layerwise probing across all five architectures and on Mamba-specific Conv1D attribution, so the grouping tracks computational structure rather than commutativity. Causal interventions show that, in the 4-layer formal models, linearly readable directions are often functionally necessary and can remain important at out-of-distribution lengths on Parity and Dyck. At pretrained scale the picture splits. Fine-tuned Pythia Dyck has a strong middle-layer bottleneck (L6-L7 ablation drops accuracy by roughly 81% at 160M; broader L4-L18 plateau at 410M), far weaker at the best-probe layer. Pretrained Mamba shows the complementary failure mode: its final layer is highly readable, no single probe direction breaks the task on Parity, Dyck, or S_3, yet mid-position activation patching there recovers about 97-98% of the clean-corrupted logit gap. Probing localizes where state is linearly available, not always where the computation is bottlenecked. Mechanistic signatures are properties of architecture and task together.
- Abstract(参考訳): シーケンスモデルの力学的研究は、しばしば階層的な状態符号化をアーキテクチャ上の特性として扱う: 繰り返しモデルは可読性に集中し、注意に基づくモデルはそれを分散する。
タスクが変更されたとき、同じアーキテクチャがこのプロファイルを反転させるのが分かります。
Transformers, Mamba, Mamba-2, LSTMs, GRUsの他、ParityはMambaとリカレントベースラインの後半に集中し、Transformerによって徐々に構築される。
同じフリップが微調整されたMamba-130MとPythia-160Mに現れ、Pythia Dyckのボトルネックは410Mで持続する。
代数構造(可換性)と計算構造(事前更新対スタック)である。
それらを分離するために、第3のタスク、非可換 S_3 置換合成を加えます。
ダイクではなくパリティを持つS_3群は5つのアーキテクチャすべてとマンバ固有のConv1D属性を階層的に探索するので、群化は可換性ではなく計算構造を追跡する。
因果的介入は、4層形式モデルにおいて、線形可読方向はしばしば機能的に必要であり、パリティとディックの分布外距離において重要であることを示している。
事前訓練されたスケールでは、絵は分割されます。
微細調整されたPythia Dyckは強い中間層ボトルネックを持ち(L6-L7のアブレーションは160Mで約81%の精度で減少し、L4-L18高原は410Mで、最高のプローブ層でははるかに弱い。
最終層は可読性が高く、単一のプローブ方向がパリティ、ディック、またはS_3のタスクを壊すことはないが、中間位置の活性化パッチはクリーンに破損したロジットギャップの97-98%を回復する。
状態が線形に利用可能な場所をローカライズする。
メカニスティックシグネチャは、アーキテクチャとタスクを一緒に持つ性質である。
関連論文リスト
- Diagnosing Overhead in Dispatch Operations: Cross-architecture Observatory [1.081571058570587]
AlltoAllのディスパッチは、MoEの専門家並列性の主要なボトルネックである。
ワークロードに関する2つの仮定をテストするために、DODOCOを導入します。
EPのスケーリングは、各アーキテクチャの計測可能な範囲内で、専門家ごとの最大/平均トークン比を5%以上変更する。
論文 参考訳(メタデータ) (2026-05-20T10:14:00Z) - Three Roles, One Model: Role Orchestration at Inference Time to Close the Performance Gap Between Small and Large Agents [0.4666493857924357]
複雑なマルチステップ環境において,推論時足場のみに追加のトレーニング計算を使わずに,小さなモデルの性能を向上させることができるかどうかを検討した。
我々は,AppWorldベンチマークのQwen3-8Bを,完全精度と4ビット量子化構成の両方で評価した。
本格的な推測では、私たちの足場付き8Bモデルは、オリジナルのAppWorld評価からDeepSeek-Coder 33Bインストラクション(7.1%)を上回っています。
論文 参考訳(メタデータ) (2026-04-13T13:40:33Z) - Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study [0.0]
AMD Instinct MI325X GPUにおけるLCM推定のクロスアーキテクチャ評価
3つのアーキテクチャファミリにまたがる235Bから1兆のパラメータにまたがる4つのモデルのベンチマーク。
論文 参考訳(メタデータ) (2026-02-27T13:21:48Z) - How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding [3.8914132324834045]
CoT(Chain-of- Thought)は、多段階タスクにおけるLarge Language Modelsの精度を高める。
しかし、生成された「考え」が真の内部推論過程を反映しているかどうかは未解決である。
本研究は,CoT忠実度に関する最初の特徴レベル因果研究である。
論文 参考訳(メタデータ) (2025-07-24T10:25:46Z) - Exploring Diffusion Transformer Designs via Grafting [82.91123758506876]
計算予算の少ない新しいアーキテクチャを実現するために,事前に訓練された拡散変換器(DiT)を編集する簡単な手法であるグラフト方式を提案する。
演算子置換からアーキテクチャ再構成に至るまで,事前訓練したDiTをグラフトすることで,新しい拡散モデルの設計を探索できることが示されている。
論文 参考訳(メタデータ) (2025-06-05T17:59:40Z) - Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners [72.37408197157453]
近年の進歩により、大規模言語モデル(LLM)の性能は、テスト時に計算資源をスケーリングすることで大幅に向上することが示されている。
複雑性が低いモデルは、より優れた生成スループットを活用して、固定された計算予算のために同様の大きさのトランスフォーマーを上回りますか?
この問題に対処し、強い四分法的推論器の欠如を克服するために、事前訓練された変換器から純およびハイブリッドのマンバモデルを蒸留する。
論文 参考訳(メタデータ) (2025-02-27T18:08:16Z) - Mixture-of-Mamba: Enhancing Multi-Modal State-Space Models with Modality-Aware Sparsity [56.0251572416922]
状態空間モデル(SSM)は、シーケンシャルモデリングのためのトランスフォーマーの効率的な代替手段として登場した。
本稿では,Mambaブロックのモダリティ特異的パラメータ化により,モダリティを意識した疎結合を実現する新しいSSMアーキテクチャを提案する。
マルチモーダル事前学習環境におけるMixture-of-Mambaの評価を行った。
論文 参考訳(メタデータ) (2025-01-27T18:35:05Z) - The Mamba in the Llama: Distilling and Accelerating Hybrid Models [76.64055251296548]
注目層からの線形射影重みを学術的なGPU資源で再利用することにより,大規模な変換器を線形RNNに蒸留する方法を示す。
結果として得られたハイブリッドモデルは、チャットベンチマークのオリジナルのTransformerに匹敵するパフォーマンスを達成する。
また,Mambaとハイブリッドモデルの推論速度を高速化するハードウェア対応投機的復号アルゴリズムを導入する。
論文 参考訳(メタデータ) (2024-08-27T17:56:11Z) - Efficient Content-Based Sparse Attention with Routing Transformers [34.83683983648021]
自己注意は、シーケンス長に関する二次計算とメモリ要求に悩まされる。
本研究は,関心の問合せとは無関係なコンテンツへのアロケートやメモリの参加を避けるために,動的スパースアテンションパターンを学習することを提案する。
論文 参考訳(メタデータ) (2020-03-12T19:50:14Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。