論文の概要: Activation Sparsity with Weight Approximation for Faster LLM Decoding on Offloaded Weights
- arxiv url: http://arxiv.org/abs/2610.02598v1
- Date: Thu, 01 Oct 2026 23:50:52 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-10-06 00:14:30.128184
- Title: Activation Sparsity with Weight Approximation for Faster LLM Decoding on Offloaded Weights
- Title(参考訳): 重み付き重み付き高速LLM復号のための重み付き活性化空間
- Abstract要約: 我々は,アクティベーション空間を利用する場合,モデル品質と復号化性能のトレードオフを改善する。
SpAxは0に最も近いアクティベーションに関連するウェイトをスキップし、より小さなマグニチュードアクティベーションのための近似ウェイトを読み、最大のマグニチュードアクティベーションのためのオリジナルのウェイトを読み取る。
重みがCPUメモリにオフロードされているため、SpAxは16ビット重みの3.86X(最大5.57X)、2.06X(最大2.74X)、4ビット重みの2.06X(最大2.74X)でデコーディングを高速化する。
- 参考スコア(独自算出の注目度): 2.8156731197882614
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Deploying LLMs on consumer-grade GPUs with insufficient memory to hold their weights can result in prohibitively slow inference, because decoding repeatedly transfers offloaded weights from system RAM or flash storage into GPU at much lower bandwidth than local GPU-memory access. Activation sparsity reduces these transfers by skipping weights associated with zero or near-zero activations. However, as more activation contributions are omitted, model quality eventually degrades rapidly, indicating that weights associated with small-magnitude activations collectively influence model quality sharply. In this work, we improve the trade-off between model quality and decoding performance when exploiting activation sparsity. Our key idea is to replace the binary choice of whether or not to read a weight with three options: fully retain it, approximate it using a compressed weight representation, or omit it entirely. SpAx skips weights associated with activations closest to zero, reads approximate weights for smaller-magnitude activations, and reads original weights for the largest-magnitude activations. Smaller-magnitude activations attenuate the errors introduced by approximate weights, while compressed weight representations require fewer bytes to be transferred. With weights offloaded to CPU memory, SpAx speeds up decoding by 3.86X on average (up to 5.57X) with 16-bit weights and 2.06X (up to 2.74X) with 4-bit weights, at a WikiText-2 perplexity increase of at most 10%. With weights offloaded to flash storage, the speedups are 3.31X on average (up to 4.81X) and 1.54X (up to 2.03X).
- Abstract(参考訳): メモリが不足したコンシューマグレードのGPUにLLMをデプロイすると、ローカルのGPUメモリアクセスよりもはるかに低い帯域で、システムRAMやフラッシュストレージからのオフロード重量をGPUに繰り返し転送するため、推論が著しく遅くなる可能性がある。
活性化間隔は、ゼロまたはほぼゼロの活性化に関連する重みをスキップすることでこれらの移動を減少させる。
しかし、より多くの活性化コントリビューションが省略されるにつれて、モデル品質は最終的に急速に低下し、小さなマグニチュードアクティベーションに関連する重みがモデル品質に急激な影響を及ぼすことを示す。
本研究では,アクティベーション空間を利用する場合のモデル品質とデコード性能のトレードオフを改善する。
私たちのキーとなる考え方は、ウェイトを読み取るかどうかという二項選択を、3つのオプションで置き換えることです。
SpAxは0に最も近いアクティベーションに関連するウェイトをスキップし、より小さなマグニチュードアクティベーションのための近似ウェイトを読み、最大のマグニチュードアクティベーションのためのオリジナルのウェイトを読み取る。
より小さなマグニチュードアクティベーションは、近似重みによる誤差を減らし、圧縮された重み表現は転送されるバイトを少なくする。
重みがCPUメモリにオフロードされているため、SpAxは16ビット重みの3.86X(最大5.57X)、2.06X(最大2.74X)、4ビット重みの2.06X(最大2.74X)でデコーディングを高速化する。
フラッシュストレージにオフロードされた重量では、スピードアップは平均3.31倍(最大4.81倍)と1.54倍(最大2.03倍)である。
関連論文リスト
- UltraSketchLLM: Saliency-Driven Sketching for Ultra-Low Bit LLM Compression [2.4253566132113877]
UltraSketchLLMはインデックスのないスケッチベースのフレームワークで、モデル性能を維持しながら超低ビット圧縮(重量あたり0.5ビットまで)を実現する。
提案手法では,小重量の相対誤差を最小限に抑えるために,AbsMaxMinスケッチを最小にするため,重み付けを優先するための重要空間割り当て,圧縮を意識した微調整のためのストレートスルー推定器を組み込んだ。
論文 参考訳(メタデータ) (2025-06-08T16:55:42Z) - SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models [61.474101404805545]
拡散モデルは高品質なイメージを生成することができるが、スケールするにつれて、メモリ要求が増加し、より高いレイテンシがデプロイメント上の課題を引き起こす。
この制限を克服する新しい4ビット量子化パラダイムであるSVDQuantを提案する。
We reduce the memory usage for the 12B FLUX.1 models by 3.5$times$, achieved 3.0$times$ speedup over the 4-bit weight-only Quantization (W4A16) baseline。
論文 参考訳(メタデータ) (2024-11-07T18:59:58Z) - Endor: Hardware-Friendly Sparse Format for Offloaded LLM Inference [47.043257902725294]
本研究では, 圧縮率が高く, 減圧オーバーヘッドの少ない非ゼロ値に対して, 刈り取られたLLM重みの非構造スパースパターンを圧縮する新しいスパース形式を提案する。
一般的なHugingface Accelerateを使ったオフロード推論と比較して、EndorはOPT-66Bを1.70倍、Llama2-70Bを1.78倍加速する。
論文 参考訳(メタデータ) (2024-06-17T15:55:08Z) - AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration [54.692405042065815]
LLM低ビット量のみの量子化のためのハードウェアフレンドリーなアプローチであるActivation-Aware Weight Quantization (AWQ)を提案する。
AWQ は 1% の正重みしか保護せず,命令調整型 LM とマルチモーダル LM の量子化性能に優れる。
また,4ビットオンデバイスLLM/VLMに適した,効率的なフレキシブルな推論フレームワークであるTinyChatを実装した。
論文 参考訳(メタデータ) (2023-06-01T17:59:10Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。