論文の概要: DistillPath: An Efficient 22M Distilled Pathology Encoder Approaching Large Foundation Model Performance
- arxiv url: http://arxiv.org/abs/2608.17872v1
- Date: Tue, 18 Aug 2026 15:06:35 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-19 21:40:53.372626
- Title: DistillPath: An Efficient 22M Distilled Pathology Encoder Approaching Large Foundation Model Performance
- Title(参考訳): DistillPath: 大規模ファンデーションモデル性能にアプローチした2200万蒸留型エンコーダ
- Abstract要約: In this present DistillPath-KS16, from the existing 22M kaiko ViT-S/16 encoder and improveing from released pathology encoder used as frozen teachers。
レシピは教師の最終クラスのみを読み、トークンをパッチし、6000の公開スライドで訓練する。
- 参考スコア(独自算出の注目度): 12.934094486594633
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Many high-performing pathology tile encoders are now foundation models with hundreds of millions to over a billion parameters. Encoding and storing the thousands of tiles in each whole-slide image with such models is costly on commodity hardware, so compact encoders that retain useful downstream performance are a valuable alternative. We present DistillPath-KS16, which starts from the existing 22M kaiko ViT-S/16 encoder and improves it by distilling from released pathology encoders used as frozen teachers. The recipe reads only the teachers' final class and patch tokens and trains on 6,000 public slides, needing neither their DINO nor iBOT pretraining heads nor a billion-tile corpus, so it applies to any released encoder that exposes backbone tokens. We distill four teachers spanning 86M to 1.1B parameters into the same student. Every variant improves the kaiko baseline on all three benchmarks we use, EVA, HEST, and PLISM, and the strongest teacher is task-dependent. On the seven-task EVA mean, DistillPath-KS16-Virchow2 reaches $0.795$, within $0.015$ points of Virchow2, the top-scoring model in our evaluation, at about $29\times$ fewer parameters; it also scores above H0-mini and GPFM on this aggregate metric, though that advantage is task-concentrated rather than uniform. Because it remains a 22M ViT-S/16 with 384-dimensional features, DistillPath-KS16 runs more than $25\times$ faster than Virchow2. Code is available at https://github.com/RamonKaspar/DistillPath, and released model weights are available at https://huggingface.co/collections/RamonK/distillpath.
- Abstract(参考訳): 多くの高性能な病理タイルエンコーダは、今や数億から10億のパラメータを持つ基礎モデルとなっている。
このようなモデルで、各スライダー画像に数千のタイルをエンコードし、保存することは、コモディティなハードウェアでコストがかかるため、下流で有用なパフォーマンスを保持するコンパクトエンコーダは、貴重な代替手段である。
In this present DistillPath-KS16, from the existing 22M kaiko ViT-S/16 encoder and improveing from released pathology encoder used as frozen teachers。
このレシピは教師の最終クラスのみを読み出し、6,000の公開スライドでトークンをパッチし、DINOやiBOTの事前訓練や10億個のタイルコーパスを必要としないため、バックボーントークンを公開する任意のリリースエンコーダに適用できる。
86Mから1.1Bのパラメータにまたがる4人の教師を同じ学生に蒸留する。
どの亜種も、私たちが使用しているEVA、HEST、PLISMの3つのベンチマークのカイコベースラインを改善し、最も強力な教師はタスク依存である。
7タスクのEVA平均では、DistillPath-KS16-Virchow2 は、Virchhow2 の0.015$ポイント以内の$0.795$に達し、パラメータが約29\times$少なくなる。
DistillPath-KS16は2200万ViT-S/16で、384の次元を持つため、Virchow2より25ドル以上速く走る。
コードはhttps://github.com/RamonKaspar/DistillPathで利用可能で、モデルウェイトのリリースはhttps://huggingface.co/collections/RamonK/distillpathで見ることができる。
関連論文リスト
- Cross-Architecture Knowledge Distillation from a Vision Foundation Model to a Lightweight Visual State Space Model for Tea Leaf Disease Classification [10.122947222129108]
DINOv2のような自己監督型ビジョンファウンデーションモデルは強力な機能を提供するが、フィールド展開には大きすぎる。
微調整DINOv2教師からコンパクトな双方向視覚空間モデルへのクロスアーキテクチャ知識の蒸留について検討した。
我々は,SSM学生が限られたデータで学習することを防ぐ2つのトレーニング安定問題を同定し,修正する。
論文 参考訳(メタデータ) (2026-08-27T08:02:05Z) - Multi-Teacher Contrastive Distillation for Edge-Efficient Pathology Foundation Models [1.9573380763700714]
複数のPFMから小型エッジ指向エンコーダに冷凍タイル埋め込みを蒸留する事前学習フレームワーク MuCoDi を提案する。
11.8KのWSIから14.3MのTCGAタイルをプレトレーニングし、23の下流分類タスクで凍結エンコーダの評価を行った。
論文 参考訳(メタデータ) (2026-07-06T18:14:37Z) - TESSERA v2: Scaling Pixel-wise Earth Foundation Models [0.0]
本稿では,地球観測画素単位の基礎モデルについて,これまでで最大規模で評価したスケーリングモデルについて述べる。
395のトレーニングは1024GH200スーパーチップ上で、固定画素ワイドのバーロウツインズファミリー内で実行される。
埋め込み・データ配置のためのコンパクトな学生に画素単位のモデルを蒸留する。
論文 参考訳(メタデータ) (2026-07-04T16:52:34Z) - From Patches to Trajectories: Privileged Process Supervision for Software-Engineering Agents [56.31499185764872]
教師の長い軌道上の監督された微調整(SFT)は、オープンソフトウェアエンジニアリング(SWE)エージェントに調査と推論を浸透させる主要な方法である。
本稿では,P2T (Patches-to-Trajectories) を提案する。P2T (Patches-to-Trajectories) は,P2T (Patches-to-Trajectories) において,P2T (Patches-to-Trajectories) とP2T (Patches-to-Trajectories) の2つの最適化法である。
論文 参考訳(メタデータ) (2026-05-21T04:54:55Z) - RiT: Vanilla Diffusion Transformers Suffice in Representation Space [12.711808725422108]
x$prediction とのフローマッチングは、ピクセル空間 citeli2025back において、低次元多様体構造を効果的に活用することが知られている。
事前学習された表現空間は、本質的な次元に匹敵する低次元データ多様体を含むが、フローマッチング学習に好適な分布を提供するかどうかを問う。
論文 参考訳(メタデータ) (2026-05-21T04:21:43Z) - HEED: Density-Weighted Residual Alignment for Hybrid Vision-Language Model Distillation [5.501291336853232]
我々はQwen3-VL-8B-インストラクトを3:1のマンバ-2/アテンションハイブリッドに蒸留する。
学生モデルは、MMStar、MMBench、MMMU-Proといったビジュアル推論ベンチマークで教師の2ポイント以内に留まる。
学生は依然としてシーンを理解できるが、答えるために必要な細かい文章は失われる。
通常のポストトレーニングの後、学生は10ベンチマーク平均で4.12$times$スループットで教師レベルのパフォーマンスに達し、128kコンテキストで68%のメモリ節約を行う。
論文 参考訳(メタデータ) (2026-05-16T17:33:24Z) - ZenBrain: A Neuroscience-Inspired 7-Layer Memory Architecture for Autonomous AI Systems [51.56484100374058]
LongMemEval-500では、ZenBrainは長いコンテキストのオラクルのバイナリ・ジャッジの精度を4.5pp以内と一致させる。
ZenBrainは7層の神経科学にインスパイアされたメモリアーキテクチャである。
論文 参考訳(メタデータ) (2026-04-26T20:39:19Z) - SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations [54.303301888915406]
混合エキスパートモデル(MoE)は、計算コストを大幅に増加させることなく、言語モデルをスケールアップするためのデファクトアーキテクチャとして登場した。
最小のアクティベーションキャッシングでMoEの前後パスを計算するメモリ効率のアルゴリズムを提案する。
また,グループ化されたGEMMカーネルのパディングによる無駄計算を最小限に抑える新しい「トークンラウンドリング」手法を提案する。
論文 参考訳(メタデータ) (2025-12-16T04:39:10Z) - Seq vs Seq: An Open Suite of Paired Encoders and Decoders [37.62535961965971]
我々は,1700万のパラメータから10億までの,ペア付きエンコーダのみとデコーダのみのモデルであるSOTAオープンデータEttinスイートを紹介する。
エンコーダのみのモデルとデコーダのみのモデルの両方で同じレシピを使用して、それぞれのサイズで両方のカテゴリでSOTAレシピを生成する。
本稿では,デコーダモデルをエンコーダのタスク(およびその逆も)に適応させることが,逆の目的のみを使用する場合に比べて低いことを示す。
論文 参考訳(メタデータ) (2025-07-15T15:31:51Z) - Simple ReFlow: Improved Techniques for Fast Flow Models [68.32300636049008]
拡散および流れマッチングモデルは、優れた生成性能を実現するが、多くのサンプリングステップを犠牲にしている。
我々は、力学、学習、推論のトレーニングに7つの改善点を提案する。
我々は、ニューラルネットワークによる高速な生成のために、最先端のFIDスコア(ガイダンスなし/参照なし)を達成している。
論文 参考訳(メタデータ) (2024-10-10T11:00:55Z) - SdAE: Self-distillated Masked Autoencoder [95.3684955370897]
本稿では,自己蒸留マスク付きオートエンコーダネットワークSdAEを提案する。
300エポックの事前トレーニングで、バニラViT-BaseモデルはImageNet-1k分類において84.1%の微調整精度を達成する。
論文 参考訳(メタデータ) (2022-07-31T15:07:25Z) - Compressing 1D Time-Channel Separable Convolutions using Sparse Random
Ternary Matrices [65.4388266814055]
1次元時間チャネル分離可能な畳み込みの1x1-畳み込みを、定数でスパースな乱数三元行列で-1,0,+1$の重みで置き換える。
Google Speech Commands v1のコマンド認識のために、最新の精度を同じネットワークサイズで97.21%$から97.41%$に改善します。
librispeech上での音声認識では、トレーニングすべき重みの数は半分になり、浮動小数点ベースラインの単語誤り率の約1%を犠牲にします。
論文 参考訳(メタデータ) (2021-03-31T15:09:20Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。