論文の概要: Reference-Based Distillation Detection in LLMs
- arxiv url: http://arxiv.org/abs/2607.09692v1
- Date: Fri, 19 Jun 2026 06:59:19 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-19 21:54:20.37165
- Title: Reference-Based Distillation Detection in LLMs
- Title(参考訳): LLMにおける基準ベース蒸留検出
- Authors: Rajat Rawat, Sizhe Chen, Akshay Anand, Michael Duan, Bob Rotsted, Sewon Min,
- Abstract要約: 本稿では,参照型メンバシップ推論に基づく蒸留検出手法を提案する。
制御された蒸留実験と実世界のモデルの両方にまたがるハイブリッド評価を開発した。
この手法を現代モデルに適用すると、QwQ、DeepSeek-R1、GPT-OSSを含む蒸留関係に関する新たな証拠が得られる。
- 参考スコア(独自算出の注目度): 17.877444310396758
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Model distillation -- training on outputs from stronger third-party models -- is widely used to boost performance, but raises concerns about unfair advantages and policy violations. This motivates a fundamental question: can we detect whether a model was distilled from another? We show that, while identifying a teacher model from a student in isolation is highly challenging, it becomes tractable in a reference-based setting: given a model and an earlier-generation checkpoint from the same lineage, we can identify the teacher model used to train the later checkpoint. We introduce a distillation detection method based on reference-based membership inference. By comparing how strongly a student model preferentially aligns with outputs from different candidate teachers relative to a reference checkpoint, our method identifies the most likely teacher and detects evidence of distillation. To handle unknown distillation pipelines such as hidden prompts, we infer proxy prompt templates directly from model outputs. We additionally identify a distinctive glyph-level signal specific to o1/o3 models. Evaluating distillation detection is challenging because modern model lineages are already heavily entangled. To address this, we develop a hybrid evaluation spanning both controlled distillation experiments and real-world models. Across both settings, our approach recovers the true teacher with near-perfect accuracy in single-teacher distillation scenarios, even when the underlying distillation pipeline is largely unknown. We further introduce statistical tests for both teacher attribution and distillation detection, and extend our framework to open-world settings where no teacher is guaranteed to be present among the candidates. Applying our method to contemporary models yields new evidence regarding potential distillation relationships involving QwQ, DeepSeek-R1, and GPT-OSS.
- Abstract(参考訳): より強力なサードパーティモデルの出力をトレーニングするモデル蒸留は、パフォーマンス向上に広く利用されているが、不当な優位性や政策違反に対する懸念が高まっている。
モデルが別のモデルから蒸留されたかどうかを検出できますか?
そこで本研究では,教師モデルと生徒を分離して識別することは極めて困難であるが,参照ベースの設定では,同じ系統のモデルと初期チェックポイントが与えられた場合,後者のチェックポイントのトレーニングに使用される教師モデルを特定することができることを示す。
本稿では,参照型メンバシップ推論に基づく蒸留検出手法を提案する。
本手法は, 生徒モデルが, 基準チェックポイントに対して, 学生モデルと異なる教師の出力とを優先的に比較することにより, 最も可能性の高い教師を識別し, 蒸留の証拠を検出する。
隠れプロンプトなどの未知の蒸留パイプラインを扱うために、モデル出力から直接プロキシプロンプトテンプレートを推論する。
また、o1/o3モデル特有のグリフレベルの信号も同定する。
現代のモデル系統は、既に強く絡み合っているため、蒸留検出の評価は困難である。
そこで我々は,制御された蒸留実験と実世界のモデルの両方にまたがるハイブリッド評価を開発した。
いずれの設定においても,本手法は,基礎となる蒸留パイプラインがほとんど不明な場合においても,単教師蒸留シナリオにおける真の教師の精度をほぼ完全に回復する。
さらに,教師の帰属と蒸留検知の両面での統計的テストを導入し,その枠組みを,候補者に教師がいないようなオープンワールド環境に拡張する。
この手法を現代モデルに適用すると、QwQ、DeepSeek-R1、GPT-OSSを含む蒸留関係に関する新たな証拠が得られる。
関連論文リスト
- Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation [44.23725347095524]
クロスモデル推論蒸留防止トラクションフレームワークについて紹介する。
蒸留モデルにより生成された各行動について,教師,元学生,蒸留モデルに割り当てられた予測確率を同じ文脈で求める。
実験により, 蒸留モデルでは, 実際に教師が選択した行動が生成され, 実験結果と相関し, 測定結果が妥当に説明できることが実証された。
論文 参考訳(メタデータ) (2025-12-24T03:19:05Z) - Warmup-Distill: Bridge the Distribution Mismatch between Teacher and Student before Knowledge Distillation [84.38105530043741]
本稿では, 学生の蒸留を教員の蒸留と整合させて, 蒸留に先立って行うワームアップ蒸留法を提案する。
7つのベンチマークの実験は、ウォームアップ・ディスティルが蒸留に適したウォームアップの学生を提供することを示した。
論文 参考訳(メタデータ) (2025-02-17T12:58:12Z) - Towards Training One-Step Diffusion Models Without Distillation [72.80423908458772]
我々は,教師のスコア管理を完全に禁止する,新しい研修方法のファミリーを紹介する。
教師の重みによる学生モデルの初期化は依然として重要な課題である。
論文 参考訳(メタデータ) (2025-02-11T23:02:14Z) - Knowledge Distillation with Refined Logits [31.205248790623703]
本稿では,現在のロジット蒸留法の限界に対処するため,Refined Logit Distillation (RLD)を導入する。
我々のアプローチは、高性能な教師モデルでさえ誤った予測をすることができるという観察に動機づけられている。
本手法は,教師からの誤解を招く情報を,重要なクラス相関を保ちながら効果的に排除することができる。
論文 参考訳(メタデータ) (2024-08-14T17:59:32Z) - Unbiased Knowledge Distillation for Recommendation [66.82575287129728]
知識蒸留(KD)は推論遅延を低減するためにレコメンダシステム(RS)に応用されている。
従来のソリューションは、まずトレーニングデータから完全な教師モデルを訓練し、その後、その知識を変換して、コンパクトな学生モデルの学習を監督する。
このような標準的な蒸留パラダイムは深刻なバイアス問題を引き起こし、蒸留後に人気アイテムがより強く推奨されることになる。
論文 参考訳(メタデータ) (2022-11-27T05:14:03Z) - Why distillation helps: a statistical perspective [69.90148901064747]
知識蒸留は、単純な「学生」モデルの性能を向上させる技術である。
この単純なアプローチは広く有効であることが証明されているが、基本的な問題は未解決のままである。
蒸留が既存の負の鉱業技術をどのように補完し, 極端に多層的検索を行うかを示す。
論文 参考訳(メタデータ) (2020-05-21T01:49:51Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。