論文の概要: Progressive$^2$: A Teacher-Student Progressive Co-Evolving Knowledge Distillation Method for Substantial Model Compression
- arxiv url: http://arxiv.org/abs/2608.00129v1
- Date: Fri, 31 Jul 2026 13:19:08 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-04 15:07:24.485525
- Title: Progressive$^2$: A Teacher-Student Progressive Co-Evolving Knowledge Distillation Method for Substantial Model Compression
- Title(参考訳): プログレッシブ$^2$: 実体モデル圧縮のための教師と学習者の共進化的知識蒸留法
- Authors: Tiancong Cheng, Ying Zhang, Zhiwen Yu, Yifang Yin, Bin Guo,
- Abstract要約: 知識蒸留(KD)は、大きなモデル(教師)からより小さなモデル(学生)へ知識を伝達するための広く利用されている技術である。
本稿では, より強力な教師と, より小さな学生のコンビネーションを通じて, プログレッシブ$2$という新しい蒸留手法を提案する。
- 参考スコア(独自算出の注目度): 19.122713439477273
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Knowledge distillation (KD) is a widely utilized technique for transferring knowledge from a large model (the teacher) to a smaller model (the student). Owing to its flexibility and broad applicability, KD has been extensively applied in the compression of server-side models to meet the Quality of Service (QoS) requirements of client users. Despite significant advancements, the performance of distillation is substantially compromised when a large disparity exists between the capabilities of the server and the requirements of the client. To alleviate this problem, we propose a novel distillation approach, named Progressive$^2$, which operates through the combination of a progressively stronger teacher and a progressively smaller student. On the side of the teacher, rather than involving all layers simultaneously, we progressively select additional layers for distillation following a raw-to-rich semantic progression, establishing a systematic learning curriculum. Furthermore, we design a teacher-side multi-feature fusion adapter for the teacher to improve training stability, which is theoretically supported by the framework of Lipschitz continuity. On the side of the student, rather than directly training a tiny model, we gradually reduce the size of the network to facilitate an iterative co-evolution with the teacher. Progressive$^2$ serves as a flexible framework; the progressive strategy of the teacher can be deployed independently to achieve an optimal balance between accuracy and training efficiency, while the joint integration of the teacher and the student yields further improvements in overall performance.
- Abstract(参考訳): 知識蒸留(KD)は、大きなモデル(教師)からより小さなモデル(学生)へ知識を伝達する技術として広く利用されている。
柔軟性と広範な適用性のため、KDはクライアントユーザのQuality of Service(QoS)要件を満たすために、サーバサイドモデルの圧縮に広く適用されています。
大幅な進歩にもかかわらず、サーバの能力とクライアントの要求との間に大きな相違が存在する場合、蒸留の性能は著しく損なわれる。
この問題を軽減するために,プログレッシブ$^2$という新しい蒸留手法を提案する。
教師の側では,全ての層を同時に巻き込むのではなく,生から豊かなセマンティック・プログレクションの後に,蒸留のための追加層を段階的に選択し,体系的な学習カリキュラムを確立する。
さらに,リプシッツ連続性の枠組みを理論的に支持する教師側多機能核融合アダプタを設計し,訓練安定性を向上させる。
学生側では、小さなモデルを直接訓練するのではなく、ネットワークのサイズを徐々に小さくし、教師との反復的共進化を促進する。
プログレッシブ$^2$はフレキシブルなフレームワークとして機能し、教師のプログレッシブ戦略を独立して展開して精度とトレーニング効率の最適なバランスを保ちながら、教師と学生の連携によって全体的なパフォーマンスがさらに向上する。
関連論文リスト
- AdaSwitch: Adaptive Switching Generation for Knowledge Distillation [58.647880811071495]
スモール言語モデル(SLM)は、厳密な待ち時間と計算制約のあるアプリケーションには不可欠である。
トークンレベルでのオン・ポリティクスとオフ・ポリティクス・ジェネレーションを組み合わせた新しいアプローチであるAdaSwitchを提案する。
AdaSwitchは一貫して精度を向上し、SLMを蒸留するための実用的で効果的な方法を提供し、追加のオーバーヘッドを許容する。
論文 参考訳(メタデータ) (2025-10-09T06:38:37Z) - Progressive Weight Loading: Accelerating Initial Inference and Gradually Boosting Performance on Resource-Constrained Environments [8.020686883632594]
プログレッシブウェイトローディング(Progressive Weight Loading, PWL)は、最初は軽量の学生モデルをデプロイし、次にその層を事前訓練された教師モデルに置き換えることで、高速な初期推論を可能にする技術である。
VGG, ResNet, ViT アーキテクチャに関する実験により,PWL で訓練されたモデルは,教師層がロードされるにつれて,競争蒸留性能を維持し,徐々に精度を向上することを示した。
論文 参考訳(メタデータ) (2025-09-26T13:19:32Z) - Intra-class Patch Swap for Self-Distillation [3.282914142012984]
単一学生ネットワークに基づく無教師蒸留フレームワークを提案する。
我々のアプローチは、クラス内パッチスワップスワップスワップスワップスワップスワップスワップスワップスワップスワップスワップスワップスワップスワップスワップスワップスワップスワップスワップスワップスワップスワップスワップスワップスワ
提案手法は,既存の自己蒸留ベースラインと従来の教師ベースのKDアプローチを一貫して上回る。
論文 参考訳(メタデータ) (2025-05-20T09:30:19Z) - Exploring and Enhancing the Transfer of Distribution in Knowledge Distillation for Autoregressive Language Models [62.5501109475725]
知識蒸留(KD)は、より小さな学生モデルを模倣するように訓練することで、大きな教師モデルを圧縮する技術である。
本稿では、教師ネットワークが小さなオンラインモジュールを統合し、学生モデルと同時学習するオンライン知識蒸留(OKD)について紹介する。
OKDは、様々なモデルアーキテクチャやサイズにおけるリードメソッドのパフォーマンスを達成または超え、トレーニング時間を最大4倍に短縮する。
論文 参考訳(メタデータ) (2024-09-19T07:05:26Z) - Generalizing Teacher Networks for Effective Knowledge Distillation Across Student Architectures [4.960025399247103]
Generic Teacher Network (GTN) は、知識を有限のアーキテクチャプールからサンプリングされた任意の学生モデルに効果的に伝達できる汎用的な教師を作成するための、一発のKD-awareトレーニングである。
本手法は, 総合的なKD効果の向上と, プール内の生徒間での総合教師のトレーニングコストの最小化を両立させる。
論文 参考訳(メタデータ) (2024-07-22T20:34:00Z) - Better Teacher Better Student: Dynamic Prior Knowledge for Knowledge
Distillation [70.92135839545314]
本研究では,教師の持つ特徴の一部を,特徴蒸留前の先行知識として統合した動的事前知識(DPK)を提案する。
DPKは,教員モデルと生徒モデルのパフォーマンスを正に相関させ,より大きな教員を適用することで生徒の精度をさらに高めることができる。
論文 参考訳(メタデータ) (2022-06-13T11:52:13Z) - Learning to Teach with Student Feedback [67.41261090761834]
対話的知識蒸留 (Interactive Knowledge Distillation, IKD) は、教師が生徒のフィードバックから教えることを学ぶことを可能にする。
IKDは教師モデルを訓練し、特定の学生のトレーニングステップごとに特定のソフトターゲットを生成する。
教師と生徒の協調的な最適化は2つの反復的なステップによって達成される。
論文 参考訳(メタデータ) (2021-09-10T03:01:01Z) - Peer Collaborative Learning for Online Knowledge Distillation [69.29602103582782]
Peer Collaborative Learningメソッドは、オンラインアンサンブルとネットワークコラボレーションを統合フレームワークに統合する。
CIFAR-10, CIFAR-100, ImageNetによる実験により, 提案手法は種々のバックボーンネットワークの一般化を著しく改善することを示した。
論文 参考訳(メタデータ) (2020-06-07T13:21:52Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。