論文の概要: Investigating Model Compression for Neural Machine Translation in the Biomedical Domain
- arxiv url: http://arxiv.org/abs/2610.07032v1
- Date: Sun, 04 Oct 2026 19:01:38 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-08 02:58:29.521531
- Title: Investigating Model Compression for Neural Machine Translation in the Biomedical Domain
- Title(参考訳): バイオメディカル領域におけるニューラルネットワーク翻訳のためのモデル圧縮の検討
- Abstract要約: バイオメディカル翻訳における知識蒸留と定量化の併用について検討した。
共同蒸留・定量化学生モデルでは, サイズが69%, 推算速度が98.21%, 二酸化炭素排出量が98.46%減少した。
- 参考スコア(独自算出の注目度): 0.5043189915779883
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Large-scale pretrained transformer models have achieved state-of-the-art performance across diverse machine translation tasks, including multilingual settings. Knowledge distillation has emerged as a sustainable approach for model compression, transferring knowledge from large teacher models to smaller, more efficient student models. Similarly, quantization, which reduces the numerical precision of model weights and activations (e.g., from 32-bit to 8-bit representations) is widely used to accelerate inference, enabling models to run several times faster during deployment. However, both techniques face limitations when applied to specialized domain data, particularly under low-resource conditions. In knowledge distillation, the effectiveness of transfer is often constrained by the scarcity of domain-specific parallel data, while quantization can lead to performance degradation as bit precision decreases. In this work, we investigate the combined application of knowledge distillation and quantization for French-to-English biomedical translation, a domain characterized by specialized terminology and limited parallel resources. We develop and compare multiple fine-tuning strategies to adapt compressed student models to this challenging setting. Our experiments demonstrate that a collaboratively distilled and quantized student model achieves a 69% reduction in size, a 98.21% increase in inference speed, and a 98.46% reduction in CO2 emissions compared to the original baseline all without sacrificing translation quality. These results indicate that jointly optimized compression techniques can yield efficient, high-performance models suitable for translation service providers operating under resource constraints.
- Abstract(参考訳): 大規模事前学習型トランスフォーマーモデルは、多言語設定を含む多種多様な機械翻訳タスクで最先端のパフォーマンスを達成した。
知識蒸留はモデル圧縮のための持続可能なアプローチとして出現し、大きな教師モデルからより小さく、より効率的な学生モデルに知識を移す。
同様に、モデルウェイトとアクティベーションの数値的精度を下げる量子化(例えば、32ビットから8ビットの表現)は、推論を加速するために広く使われ、モデルが展開中に数倍高速に動作できるようにする。
しかし、どちらの手法も、特に低リソース条件下で、特別なドメインデータに適用する場合に制限に直面している。
知識蒸留では、転送の有効性はドメイン固有の並列データの不足によって制約されることが多いが、量子化はビット精度が低下するにつれて性能劣化につながる。
本研究では, 専門用語と限られた並列資源を特徴とする領域であるフランス語と英語のバイオメディカル翻訳における知識蒸留と定量化の併用について検討する。
我々は、圧縮された学生モデルをこの困難な状況に適応させるために、複数の微調整戦略を開発し、比較する。
実験の結果, 共同蒸留・定量化学生モデルでは, サイズが69%減少し, 推算速度が98.21%向上し, 翻訳品質を損なうことなくCO2排出量が98.46%減少した。
これらの結果から, 資源制約下で動作する翻訳サービスプロバイダに適した, 効率的かつ高性能なモデルが得られることが示唆された。
関連論文リスト
- Integrating Pruning with Quantization for Efficient Deep Neural Networks Compression [0.0]
プルーニングと量子化は、モデルのサイズを減らし、処理速度を向上させるために広く使われている圧縮技術である。
本稿では,類似性に基づくフィルタプルーニングとアダプティブ・パワー・オブ・ツー(APoT)量子化を統合し,高い圧縮効率を実現する2つの手法を提案する。
実験により,提案手法は精度の低下を最小限に抑え,効率的なモデル圧縮を実現することを示す。
論文 参考訳(メタデータ) (2025-09-04T14:17:28Z) - Compression Strategies for Efficient Multimodal LLMs in Medical Contexts [0.05999777817331314]
本稿では、医療応用のための微調整LAVAモデルにおける構造解析とアクティベーション対応量子化の影響について検討する。
本研究では, プルー・SFT量子化パイプラインにおいて, 異なる量子化手法を解析し, 性能トレードオフを評価する新しい層選択法を提案する。
論文 参考訳(メタデータ) (2025-07-29T16:25:51Z) - Effective Interplay between Sparsity and Quantization: From Theory to Practice [33.697590845745815]
組み合わせると、空間性と量子化がどう相互作用するかを示す。
仮に正しい順序で適用しても、スパーシリティと量子化の複合誤差は精度を著しく損なう可能性があることを示す。
我々の発見は、資源制約の計算プラットフォームにおける大規模モデルの効率的な展開にまで及んでいる。
論文 参考訳(メタデータ) (2024-05-31T15:34:13Z) - What Happens When Small Is Made Smaller? Exploring the Impact of Compression on Small Data Pretrained Language Models [2.2871867623460216]
本稿では, AfriBERTa を用いた低リソース小データ言語モデルにおいて, プルーニング, 知識蒸留, 量子化の有効性について検討する。
実験のバッテリを用いて,圧縮が精度を超えるいくつかの指標のパフォーマンスに与える影響を評価する。
論文 参考訳(メタデータ) (2024-04-06T23:52:53Z) - QuEST: Low-bit Diffusion Model Quantization via Efficient Selective Finetuning [52.157939524815866]
本稿では,不均衡な活性化分布を量子化困難の原因として同定する。
我々は,これらの分布を,より量子化しやすいように微調整することで調整することを提案する。
本手法は3つの高解像度画像生成タスクに対して有効性を示す。
論文 参考訳(メタデータ) (2024-02-06T03:39:44Z) - Quantize Once, Train Fast: Allreduce-Compatible Compression with Provable Guarantees [53.950234267704]
我々は、全リデュース勾配互換量子化法であるGlobal-QSGDを紹介する。
ベースライン量子化法で最大3.51%の分散トレーニングを高速化することを示す。
論文 参考訳(メタデータ) (2023-05-29T21:32:15Z) - Too Brittle To Touch: Comparing the Stability of Quantization and
Distillation Towards Developing Lightweight Low-Resource MT Models [12.670354498961492]
最先端の機械翻訳モデルは、しばしば低リソース言語のデータに適応することができる。
知識蒸留(Knowledge Distillation)は、競争力のある軽量モデルを開発するための一般的な技術である。
論文 参考訳(メタデータ) (2022-10-27T05:30:13Z) - What Do Compressed Multilingual Machine Translation Models Forget? [102.50127671423752]
平均BLEUはわずかに減少するが,表現不足言語の性能は著しく低下する。
圧縮は,高リソース言語においても,本質的な性差や意味バイアスを増幅することを示した。
論文 参考訳(メタデータ) (2022-05-22T13:54:44Z) - Automatic Mixed-Precision Quantization Search of BERT [62.65905462141319]
BERTのような事前訓練された言語モデルは、様々な自然言語処理タスクにおいて顕著な効果を示している。
これらのモデルは通常、数百万のパラメータを含んでおり、リソースに制約のあるデバイスへの実践的なデプロイを妨げている。
本稿では,サブグループレベルでの量子化とプルーニングを同時に行うことができるBERT用に設計された混合精密量子化フレームワークを提案する。
論文 参考訳(メタデータ) (2021-12-30T06:32:47Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。