論文の概要: Decomposing Runtime, Kernel, and Quantization Speedups via a Matched FP16 Intermediate: A Hardware-Conditioned Case Study on Four NVIDIA RTX A5000 GPUs
- arxiv url: http://arxiv.org/abs/2607.11368v1
- Date: Mon, 13 Jul 2026 10:31:54 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-14 17:47:21.436609
- Title: Decomposing Runtime, Kernel, and Quantization Speedups via a Matched FP16 Intermediate: A Hardware-Conditioned Case Study on Four NVIDIA RTX A5000 GPUs
- Title(参考訳): マッチングされたFP16中間体による実行時,カーネル,量子化の高速化:NVIDIA RTX A5000 GPU4つのハードウェア設計事例
- Authors: Weijia Han, Lisha Qu,
- Abstract要約: 量子化は、再現可能な半精度メモリの崖を約4回通り、同時ユーザを拡張します。
一致した中間スタックは、量子化されたカーネルがフルスピードアップをランタイム部とカーネルと量子化部に分割することなく、高速なランタイムを保持する。
2つのスタック間のサンプリングモードとプロンプトプールの違いは、妥当性の脅威として記録される。
- 参考スコア(独自算出の注目度): 0.614481021961242
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Reported serving speedups from quantized kernels typically bundle the weight format, the kernel, and the inference runtime into one number. We present an attribution study on four NVIDIA RTX A5000 GPUs, 24 GiB each, on a single host with NVLink-bridged pairs. A matched intermediate stack that keeps the faster runtime without the quantized kernel splits the full speedup into a runtime part and a kernel and quantization part. Under matched greedy decoding the full stack reaches $2.58\times$ end to end, with the runtime change accounting for about two thirds of that gain on a logarithmic scale; across three similar model families the kernel and quantization part moves by at most 1.5%. Sharding one instance across all four cards falls well below doubling: a profiler trace attributes about 80% of the per token shortfall to coordination, and an NVLink versus PCIe control on the same hardware shows similar realized bandwidth on both links, pointing away from link bandwidth as the cause. Whether to run one sharded instance or several independent ones depends on the workload and the model, with the ranking reversing on the larger model: the smaller model splits between sharding and multiple instances by workload, while the larger model favors two paired instances on every workload. Quantization extends sustainable concurrent users roughly four times past a reproducible half precision memory cliff. Differences in sampling mode and prompt pool between the two stacks are documented as threats to validity.
- Abstract(参考訳): 報告された量子化されたカーネルからのサービススピードアップは、一般的に重み形式、カーネル、推論ランタイムを1つの数にまとめる。
NVLinkをブリッジした単一ホスト上で,4つのNVIDIA RTX A5000 GPU,24 GiBのコントリビューションスタディを示す。
量子化されたカーネルなしで高速なランタイムを維持するマッチした中間スタックは、フルスピードアップをランタイム部とカーネルおよび量子化部とに分割する。
一致したgreedyデコードの下では、フルスタックは2.58\times$ end to endに到達し、実行時の変更は、対数スケールでその利得の約3分の2を占める。
プロファイラトレースはトークン単位の欠点の約80%を調整に当てはめ、同じハードウェア上のNVLink対PCIeコントロールは、リンク帯域幅を原因として、両方のリンク上でも同様に実現された帯域幅を示している。
1つのシャーディングされたインスタンスを実行するか、複数の独立したインスタンスを実行するかは、ワークロードとモデルに依存する。
量子化は、再現可能な半精度メモリの崖を約4回過ぎて、持続可能な同時ユーザを拡張します。
2つのスタック間のサンプリングモードとプロンプトプールの違いは、妥当性の脅威として記録される。
関連論文リスト
- LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation [84.86814271505109]
LongLive-2.0は、長いビデオ生成の完全なトレーニングと推論ワークフロー全体を通じて、NVFP4ベースの並列インフラストラクチャである。
トレーニングには,quence-parallel autoregressive (AR) トレーニングを導入する。
実験ではトレーニングで2.15倍、推論で1.84倍のスピードアップを示す。
論文 参考訳(メタデータ) (2026-05-18T17:57:03Z) - COREY: Entropy-Guided Runtime Chunk Scheduling for Selective Scan Kernels [11.316541559874864]
プロトタイプスケジューラは、固定幅ヒストグラムを用いて推定したアクティベーションエントロピーを、チャンクサイズ選択のランタイム信号として利用する。
COREYはConcept and Feasibilityのコントリビューションとして位置づけられている。
この作業には、Tier 2aとTier 2bを接続する完全なエンドツーエンド実行が含まれていない。
論文 参考訳(メタデータ) (2026-04-12T12:07:48Z) - ParallelKittens: Systematic and Practical Simplification of Multi-GPU AI Kernels [40.94392896555992]
既存のシステムは、計算通信の重複によってこれを緩和するが、しばしばワークロードと新しいアクセラレータ間の理論的帯域幅を満たさない。
演算子固有のテクニックの代わりに、簡単な再利用可能な原則の小さなセットが、ワークロードの最適なパフォーマンスを導くことができるかどうかを問う。
PKKittens(PK)カーネルは、最大2.33倍の並列ワークロードを実現する。
論文 参考訳(メタデータ) (2025-11-17T21:48:33Z) - MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models [58.3342517278868]
本稿では,Mixed-precision AutoRegressive LINearカーネルの設計について述べる。
バッチサイズは16-32までサポートでき、量子化のスピードアップが最大 (4times$) になる。
MarLINは非同期メモリアクセス、複雑なタスクスケジューリング、パイプライン化といったテクニックを組み合わせてこれを実現している。
論文 参考訳(メタデータ) (2024-08-21T16:10:41Z) - AsyncDiff: Parallelizing Diffusion Models by Asynchronous Denoising [49.785626309848276]
AsyncDiffは、複数のデバイスにまたがるモデル並列化を可能にする、普遍的でプラグアンドプレイのアクセラレーションスキームである。
安定拡散 v2.1 では、AsyncDiff は2.7倍の速度アップと4.0倍のスピードアップを実現し、CLIPスコアの 0.38 をわずかに削減した。
我々の実験は、AsyncDiffがビデオ拡散モデルに容易に適用でき、性能を向上できることを示した。
論文 参考訳(メタデータ) (2024-06-11T03:09:37Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。