Fugu-MT 論文翻訳(概要): Comparative Characterization of KV Cache Management Strategies for LLM Inference

論文の概要: Comparative Characterization of KV Cache Management Strategies for LLM Inference

arxiv url: http://arxiv.org/abs/2604.05012v1
Date: Mon, 06 Apr 2026 16:00:39 GMT
ステータス: 翻訳完了
システム内更新日: 2026-04-08 17:42:09.405903
Title: Comparative Characterization of KV Cache Management Strategies for LLM Inference
Title（参考訳）: LLM推論のためのKVキャッシュ管理手法の比較評価
Authors: Oteo Mamo, Olga Kogiou, Hyunjin Yi, Weikuan Yu,
Abstract要約: 大言語モデル(LLM)を用いた効率的な推論にはキーバリューキャッシュが不可欠であるこれらのキャッシュは、自己回帰トークン生成時の冗長な計算を最小限にするために必須である。 KVキャッシュの成長は、システムレベルの大きな課題を引き起こしている。
参考スコア（独自算出の注目度）: 0.31498833540989407
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Abstract: Efficient inference with Large Language Models (LLMs) increasingly relies on Key-Value (KV) caches to store previously computed key and value vectors at each layer. These caches are essential to minimize redundant computation during autoregressive token generation, lowering computational complexity from quadratic to linear. However, the growth of KV caches has posed significant system-level challenges, particularly as model sizes increase, context lengths grow, and concurrent requests compete for limited memory resources. Even though several recent frameworks for KV cache management have emerged, their comparative trade-offs in memory consumption and inference performance have not been fully understood, especially under varying request sizes and model configurations. In this work, we conduct an empirical study of three state-of-the-art KV cache management frameworks: vLLM, InfiniGen, and H2O. These frameworks employ techniques such as tensor offloading, token eviction heuristics, and speculative scheduling to balance memory usage and performance. We evaluate their performance in terms of a range of metrics such as latency, throughput, and memory usage across a spectrum of key parameters including request rates, model sizes, and sparsity levels. Our results pinpoint the conditions for each framework to perform the best, revealing the most suitable selection and configuration of KV cache strategies under memory and performance constraints.
Abstract（参考訳）: LLM(Large Language Models)による効率的な推論は、以前計算されたキーと値ベクトルを各レイヤに格納するためにキーバリュー(KV)キャッシュに依存している。これらのキャッシュは、自己回帰トークン生成時の冗長な計算を最小限に抑え、計算の複雑さを2次から線形に減らすために不可欠である。しかしながら、KVキャッシュの成長は、特にモデルサイズの増加、コンテキストの長さの増加、メモリリソースの制限に対する同時要求など、システムレベルの大きな課題を引き起こしている。 KVキャッシュ管理のための最近のフレームワークがいくつか登場したが、メモリ消費と推論性能の比較トレードオフは、特に要求サイズやモデル構成の違いによって完全には理解されていない。本研究では,3つの最先端KVキャッシュ管理フレームワーク,vLLM,InfiniGen,H2Oについて実証的研究を行った。これらのフレームワークは、テンソルオフロード、トークン消去ヒューリスティックス、メモリ使用量と性能のバランスをとるための投機的スケジューリングといったテクニックを採用している。レイテンシ、スループット、メモリ使用量など、要求率、モデルサイズ、スパーシリティレベルを含む重要なパラメータの範囲で、それらのパフォーマンスを評価する。その結果,メモリおよび性能制約下でのKVキャッシュ戦略の最適選択と構成を明らかにすることで,各フレームワークが最善を尽くす条件を明らかにした。

論文の概要: Comparative Characterization of KV Cache Management Strategies for LLM Inference

関連論文リスト