論文の概要: What Does Attention Transfer Transfer? Attention Structure and Robustness in Vision Transformers
- arxiv url: http://arxiv.org/abs/2608.18399v1
- Date: Wed, 19 Aug 2026 00:10:38 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-20 20:13:55.236445
- Title: What Does Attention Transfer Transfer? Attention Structure and Robustness in Vision Transformers
- Title(参考訳): 注意伝達とは何か : 視覚変換器の注意構造とロバスト性
- Abstract要約: ビジョントランスフォーマーは、事前訓練された教師の注意マップをコピーします。
微調整の分布精度のほとんどを回復するが、分布シフト下では測定不可能に低下する。
コピーが提供するものは、アテンション構造において直接測定されることがなく、ロバスト性に結びついている。
- 参考スコア(独自算出の注目度): 0.0
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Vision transformers (ViTs) trained to copy a pretrained teacher's attention maps recover most of fine-tuning's in-distribution accuracy yet fall measurably short of it under distribution shift, as recent work has shown. What the copy delivers has never been measured directly in the attention structure and tied to robustness. We build that instrumentation for ViT-S students of a self-supervised teacher on ImageNet-100, and report three findings that triangulate one conclusion. First, the transfer is essentially perfect and permanently so: the distilled student's attention ends up roughly two orders of magnitude closer to the teacher's than fine-tuning does, and does not drift with additional training. Second, the gap is real at 14$\times$ fewer parameters and 10$\times$ less data than previously studied, but it has a time axis. It tracks training maturity, and completing the schedules that the stopping rule interrupted closes it below our pre-registered threshold in two of three seeds, with comparisons at equal accuracy giving the same result. The endpoint gap at this scale is substantially a training-maturity artifact: robustness matures later than accuracy, and stopping rules tuned to accuracy undersample it. Third, forcing cross-row redundancy down by half the structural separation between the distilled and fine-tuned conditions produces no detectable robustness response under two registered ways of matching accuracy. Verified transfer, a gap that closes while the structure never moves, and a null under direct intervention are together consistent with the deficit residing in features, not in the visible attention structure. This is elimination plus intervention, and its scope is the regime we measured. In this regime, attention overlays show where a model looks, not what it knows.
- Abstract(参考訳): 教師の注意マップをコピーするように訓練された視覚トランスフォーマー(ViTs)は、微調整のin-distriionの精度のほとんどを回復するが、最近の研究が示すように、分布シフト下では測定不可能に低下する。
コピーが提供するものは、アテンション構造において直接測定されることがなく、ロバスト性に結びついている。
我々は、ImageNet-100で教師を自力で指導するVT-S学生のために、その楽器を構築し、一つの結論を三角測量する3つの知見を報告した。
蒸留された学生の注意は、微調整よりも教師の約2桁近くまで達し、追加の訓練を伴わない。
第二に、ギャップは14$\times$より少ないパラメータと10$\times$より小さいデータであるが、時間軸を持っている。
トレーニングの成熟度を追跡し、停止規則が中断されたスケジュールが、登録済みのしきい値以下を3つのシードのうち2つで閉じる。
このスケールのエンドポイントギャップは、ほぼトレーニング成熟したアーティファクトであり、堅牢性は正確性よりも遅く成熟し、正確性に基づいて調整されたルールを停止する。
第3に、蒸留条件と微調整条件の間の構造的分離を半分にすることで、2つの登録されたマッチング精度の下では、検出可能なロバスト性応答が得られない。
確認された伝達、構造物が動かない間に閉じるギャップ、直接介入中のヌルは、目に見える注意構造ではなく特徴の欠如と一致している。
これは排除と介入であり、その範囲は私たちが測定した体制です。
この体制では、注意のオーバーレイはモデルが何を知っているかではなく、どこに見えるかを示しています。
関連論文リスト
- Learning to Track from Privileged Target Appearances [67.80976117600649]
ボトルネックを、現在のターゲット作物を正確に供給する、デプロイ不可能なオラクルと定量化する。
我々は,これらの特権的外観をデプロイ可能なトラッカーに転送する教師学生向けトレーニングフレームワークであるPrivileged Appearance Transfer for Tracking (PATT)を紹介した。
PATTは、長期追跡プロトコルと短期追跡プロトコルの両方で一貫した利得を達成する。
論文 参考訳(メタデータ) (2026-09-02T11:43:43Z) - Geometric Self-Distillation for Reasoning Generalization [12.38444431260744]
オンライン蒸留は、大規模言語モデルの訓練後の実践的なレシピである。
特権的内容の自己蒸留では、教師と学生は同じプレフィックスで条件付けられた同じモデルである。
我々は,このドリフトを学生の予測行動の運動として扱う幾何学的自己蒸留の目的であるGeoSDを紹介する。
論文 参考訳(メタデータ) (2026-07-07T23:16:19Z) - DriftScope: Measuring The Hidden Effects of Diffusion Model Adaptation [50.584834483403675]
新しい視覚概念を学習するか、望ましくないものを消去するか、事前学習したテキスト・画像拡散モデルに適応させることは、意図した効果だけで日常的に評価される。
適応が意味論的に無関係な概念を、構造的にメトリクスを集約することができない方法で体系的に損なうことを実証する。
DriftScopeは、任意の2つのモデルチェックポイントを取り込み、視覚的概念が最も移行したトークンのランクリストを返す、プロンプトレベルの診断ツールである。
論文 参考訳(メタデータ) (2026-06-30T20:57:51Z) - Perfect Demo Makes Poor Teacher: Learning Robust Alignment from Critical Motion Segments [53.85970147743299]
専門家によるデモンストレーションは、ロボット模倣学習における金の標準であると広く考えられている。
しかし、挿入、積み重ね、アライメントなどのきめ細かい操作のために、私たちは直感的な失敗モードを発見しました。
視覚言語モデルとアクションエキスパートをブリッジする,コンパクトな動的特徴であるSTAIRを導入する。
論文 参考訳(メタデータ) (2026-06-14T04:15:29Z) - Beyond Dark Knowledge: Mixup-Based Distillation for Reliable Predictions [4.672326975246762]
知識蒸留と混合はクラス境界における滑らかさの誘導に有効であることが証明されている。
彼らの相互作用は、特に学生のトレーニングでのみミキシングが適用される場合、よく理解されていない。
このミスマッチは,教師の指導信号が分散的混乱に支配されていることを示す。
論文 参考訳(メタデータ) (2026-06-10T14:59:26Z) - Phantom transitions in language model fine-tuning [0.0]
ほぼ同期の競合相手とコンテキスト上で言語モデルを微調整することは、しばしばサイレントに失敗する。
2つのファミリーにまたがる5つの変圧器アーキテクチャと5つのパラメータ範囲にまたがるこの構造について検討する。
位相遷移に類似した順序パラメータにおいて,鋭いカタパルト様ジャンプを観察する。
論文 参考訳(メタデータ) (2026-05-25T10:44:42Z) - RAIL: Region-Aware Instructive Learning for Semi-Supervised Tooth Segmentation in CBCT [8.60376207420468]
Region-Aware Instructive Learning (RAIL) は、CBCT歯のセグメンテーションのための2つのグループからなる半教師付きフレームワークである。
私たちのコードはhttps://github.com/Tournesol-Saturday/RAILで公開されます。
論文 参考訳(メタデータ) (2025-05-06T13:50:57Z) - Intensity Profile Projection: A Framework for Continuous-Time
Representation Learning for Dynamic Networks [50.2033914945157]
本稿では、連続時間動的ネットワークデータのための表現学習フレームワークIntensity Profile Projectionを提案する。
このフレームワークは3つの段階から構成される: 対の強度関数を推定し、強度再構成誤差の概念を最小化する射影を学習する。
さらに、推定軌跡の誤差を厳密に制御する推定理論を開発し、その表現がノイズに敏感な追従解析に利用できることを示す。
論文 参考訳(メタデータ) (2023-06-09T15:38:25Z) - Bidirectional Semi-supervised Dual-branch CNN for Robust 3D
Reconstruction of Stereo Endoscopic Images via Adaptive Cross and Parallel
Supervisions [15.879059896662723]
そこで本研究では,教師と生徒の両方の役割を兼ね備えた,2人の学習者との双方向学習手法を提案する。
具体的には、アダプティブ・クロス・スーパービジョン(ACS)とアダプティブ・パラレル・スーパービジョン(APS)の2つの自己スーパービジョンを紹介する。
学習知識は分岐方向(ACSにおける分散誘導)と平行方向(APSにおける分散誘導)の2方向に沿って流れている。
論文 参考訳(メタデータ) (2022-10-15T13:30:41Z) - Self-Point-Flow: Self-Supervised Scene Flow Estimation from Point Clouds
with Optimal Transport and Random Walk [59.87525177207915]
シーンフローを近似する2点雲間の対応性を確立するための自己教師型手法を開発した。
本手法は,自己教師付き学習手法の最先端性能を実現する。
論文 参考訳(メタデータ) (2021-05-18T03:12:42Z) - SeCo: Exploring Sequence Supervision for Unsupervised Representation
Learning [114.58986229852489]
本稿では,空間的,シーケンシャル,時間的観点から,シーケンスの基本的および汎用的な監視について検討する。
私たちはContrastive Learning(SeCo)という特定の形式を導き出します。
SeCoは、アクション認識、未トリムアクティビティ認識、オブジェクト追跡に関する線形プロトコルにおいて、優れた結果を示す。
論文 参考訳(メタデータ) (2020-08-03T15:51:35Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。