論文の概要: Knowledge Distillation from Large Reasoning Models to Compact Student Models: A Case Study on the John O Bryan Mathematics Competition
- arxiv url: http://arxiv.org/abs/2606.31048v1
- Date: Tue, 30 Jun 2026 02:34:45 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-01 18:27:19.044366
- Title: Knowledge Distillation from Large Reasoning Models to Compact Student Models: A Case Study on the John O Bryan Mathematics Competition
- Title(参考訳): 大規模推論モデルからコンパクト学生モデルへの知識蒸留:ジョン・オブライアン数学コンペティションを事例として
- Authors: Gaurab Baral, Aaditya Khanal, Yangyang Tao, Junxiu Zhou,
- Abstract要約: We build a Chain-of-Thought training corpus through a dual-agent framework。
このデータセットは、Apple Siliconハードウェア上でLow-Rank Adaptation (LoRA)を使用して、学生モデルを微調整するために使用される。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/publicdomain/zero/1.0/
- Abstract: This paper investigates knowledge distillation from a large reasoning model (DeepSeek-R1) to a compact student model (Qwen2.5-7B). Using historical problems from the John O'Bryan Mathematics Competition at Northern Kentucky University (2011-2025), we build a Chain-of-Thought (CoT) training corpus through a dual-agent framework. The dataset is used to fine-tune the student model with Low-Rank Adaptation (LoRA) on Apple Silicon hardware using the MLX framework. The base Qwen2.5-7B model achieves 64.67% accuracy on competition problems, while the DeepSeek-R1 teacher achieves 91.40%. An initial 1,000-iteration training run revealed severe overfitting, with validation loss reaching a minimum at iteration 200 before rising steadily. Based on this finding, we ran five independent training runs each limited to 200 iterations with varied random seeds to assess result stability. Across these five runs, the fine-tuned student model achieves a mean accuracy of 69.43% (std dev 0.17%) on the competition dataset, a 4.76 percentage-point improvement over the base model, and generalizes to 73.1% (std dev 0.18%) on the MATH-500 benchmark. We further study how response length affects answer quality across six reasoning levels (R1-R6): accuracy declines consistently from 69.43% at R1 (mean 220 words) to 41.9% at R6 (mean 31.2 words), with the two-person speed section most sensitive to token reduction. These results demonstrate that CoT distillation improves compact student models and that response length is a critical factor in mathematical reasoning quality.
- Abstract(参考訳): 本稿では,大規模な推論モデル (DeepSeek-R1) からコンパクトな学生モデル (Qwen2.5-7B) への知識蒸留について検討する。
北ケンタッキー大学(2011-2025)のジョン・オブライアン数学コンペティション(John O'Bryan Mathematicsコンペティション)の歴史的問題を利用して、我々は二重エージェントフレームワークを用いてCoT(Chain-of-Thought)トレーニングコーパスを構築した。
このデータセットは、MLXフレームワークを使用して、Apple Siliconハードウェア上でLoRA(Lo-Rank Adaptation)を使用して、学生モデルを微調整するために使用される。
Qwen2.5-7Bの基本モデルは64.67%、DeepSeek-R1の教師は91.40%である。
最初の1000回のトレーニングでは、厳重なオーバーフィッティングが明らかになり、検証の損失はイテレーション200で最小限に抑えられ、着実に上昇した。
この結果に基づいて、5つの独立したトレーニングを実行し、結果の安定性を評価するために、ランダムな種を多用した、200回のイテレーションに制限された。
これら5つの実行において、微調整された学生モデルは、競争データセットで69.43%(std dev 0.17%)、ベースモデルで4.76ポイント改善され、MATH-500ベンチマークで73.1%(std dev 0.18%)に一般化された。
さらに,R1では69.43%(平均220語)からR6では41.9%(平均31.2語)に減少し,トークン還元に最も敏感な2人の速度区間が得られた。
これらの結果は,CoT蒸留がコンパクトな学生モデルを改善すること,および応答長が数学的推論品質の重要な要因であることを証明している。
関連論文リスト
- When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models [18.966595547209824]
我々はGenDistillと組み合わせたハイブリッドKimi Delta Attention (Hybrid-KDA)アーキテクチャを提案する。
ログライクリフに基づく評価は,教師と学生のギャップを過小評価する。
論文 参考訳(メタデータ) (2026-03-27T16:16:23Z) - Evaluating Small Language Models for Front-Door Routing: A Harmonized Benchmark and Synthetic-Traffic Experiment [0.05586191108738562]
小型言語モデル(SLM)は、サブ秒、ゼロマージナルコスト、セルフホストタスクの分類に十分な推論能力を持つ。
Study 1はPhi-3.5-mini、Qwen2.5-1.5B、Qwen-2.5-3Bを同一のAzure T4ハードウェア、サービススタック、量子化、固定60ケースコーパスで同期したオフラインベンチマークである。
研究2は、合成トラフィック下で事前登録された4本腕ランダム化実験であり、有効サンプルサイズは腕あたり60ケースである。
論文 参考訳(メタデータ) (2026-03-26T15:57:46Z) - Chain of Simulation: A Dual-Mode Reasoning Framework for Large Language Models with Dynamic Problem Routing [0.0]
Chain of Simulation(CoS)は、動的に問題を特別な推論戦略にルーティングする新しいデュアルモード推論フレームワークである。
CoSは、数学的問題に対する自己整合性を伴う計算フロー、空間的推論のための表現を伴う記号的状態追跡、マルチホップ推論のためのハイブリッド事実抽出という3つの異なる推論モードを採用している。
論文 参考訳(メタデータ) (2026-02-02T21:44:01Z) - Skywork Open Reasoner 1 Technical Report [51.403686909760914]
提案するSkywork-OR1は,長期チェーン・オブ・ソート(CoT)モデルのための,効果的かつスケーラブルな強化学習(RL)実装である。
DeepSeek-R1-Distillモデルシリーズをベースとして、我々のRLアプローチは顕著なパフォーマンス向上を実現している。
我々のSkywork-OR1-32Bモデルは、AIME24とAIME25ベンチマークでDeepSeek-R1とQwen3-32Bを上回っています。
論文 参考訳(メタデータ) (2025-05-28T12:56:04Z) - Reinforcement Learning for Reasoning in Large Language Models with One Training Example [117.86853102104256]
1つのトレーニング例(1ショットRLVR)を用いた強化学習は,大規模言語モデル(LLM)の算数推論能力の向上に有効であることを示す。
1ショットRLVRにおける興味深い現象として、クロスカテゴリの一般化、自己回帰の頻度の増加、テスト性能の向上の持続などを挙げる。
論文 参考訳(メタデータ) (2025-04-29T09:24:30Z) - Benchmarking Reasoning Robustness in Large Language Models [76.79744000300363]
新規データや不完全データでは,性能が著しく低下することがわかった。
これらの結果は、厳密な論理的推論に対するリコールへの依存を浮き彫りにした。
本稿では,情報不足によって引き起こされる幻覚を利用して推論ギャップを明らかにする,Math-RoBと呼ばれる新しいベンチマークを提案する。
論文 参考訳(メタデータ) (2025-03-06T15:36:06Z) - Malware Classification from Memory Dumps Using Machine Learning, Transformers, and Large Language Models [1.038088229789127]
本研究では,異なる特徴セットとデータ構成を用いたマルウェア分類タスクにおける各種分類モデルの性能について検討する。
XGBはTop 45 Featuresで87.42%の精度を達成し、他の全てのモデルを上回った。
ディープラーニングモデルはパフォーマンスが悪く、RNNは66.71%の精度でトランスフォーマーは71.59%に達した。
論文 参考訳(メタデータ) (2025-03-04T00:24:21Z) - Common 7B Language Models Already Possess Strong Math Capabilities [61.61442513067561]
本稿では,LLaMA-2 7Bモデルと事前学習を併用したモデルが,すでに強力な数学的能力を示していることを示す。
拡張スケーリングの可能性は、公開されている数学の質問の不足によって制限されている。
論文 参考訳(メタデータ) (2024-03-07T18:00:40Z) - TACRED Revisited: A Thorough Evaluation of the TACRED Relation
Extraction Task [80.38130122127882]
TACREDはリレーショナル抽出(RE)において最も大きく、最も広く使われているクラウドソースデータセットの1つである
パフォーマンスの天井に到達したのか、改善の余地はあるのか?
ラベルエラーは絶対F1テストエラーの8%を占めており、例の50%以上を可逆化する必要がある。
論文 参考訳(メタデータ) (2020-04-30T15:07:37Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。