論文の概要: DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
- arxiv url: http://arxiv.org/abs/2609.19969v1
- Date: Thu, 17 Sep 2026 09:43:10 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-20 08:55:54.205854
- Title: DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
- Title(参考訳): DeepSeek-V4.1-Flash: KVキャッシュ圧縮の限界を押し上げる
- Abstract要約: DeepSeek-V4.1-Flashは,552Bのバックボーンパラメータを持つマルチモーダルMixture-of-Experts(MoE)モデルで,最大100万トークンのコンテキストをサポートする。
Causal-Decoder (CED) アーキテクチャでは、デコード時にトークン毎に16Bパラメータを起動するが、プリフィル時に8Bパラメータしか有効にせず、エージェントワークロードのコスト効率を大幅に向上する。
- 参考スコア(独自算出の注目度): 354.65055709450576
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Together, these compute, storage, and bandwidth demands constitute the primary bottleneck to further lowering deployment costs. To address this challenge, we introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. With its Causal Encoder-Decoder (CED) architecture, the model activates 16B parameters per token during decode but only 8B parameters during prefill, substantially improving cost efficiency for agentic workloads. To push the limits of KV cache compression, DeepSeek-V4.1-Flash combines cross-layer KV cache reuse in Compressed Sparse Attention 2 (CSA2) with FP4 KV caching. These designs reduce its global KV cache footprint (always in HBM) to 890 bytes per token, roughly 1/4 of the corresponding footprint of DeepSeek-V4-Flash. Further, through a dedicated deployment optimization known as SWA Bounded Replay, DeepSeek-V4.1-Flash reduces its persistent KV cache footprint (always on SSD or in host memory) to roughly 1/8 of that of DeepSeek-V4-Flash. Despite its much smaller KV cache footprint, the model delivers substantially better performance than the baseline. In addition, we streamline the DeepSeek-V4 architecture and introduce several efficient architectural extensions. We pretrain DeepSeek-V4.1-Flash on a multimodal corpus comprising 45T tokens and conduct comprehensive post-training, yielding strong performance across diverse text-based and multimodal agentic scenarios. Model checkpoints are available at https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash.
- Abstract(参考訳): ロングホライゾンエージェントの普及により、モデルワークロードはますますインプットヘビーになっている。
以前の作業では、長いコンテキスト計算のコストを大幅に削減したが、プリフィルは計算コストが高く、大規模なKVキャッシュはHBMとSSD容量とデータ転送帯域を圧迫し続けている。
これらの計算、ストレージ、帯域幅の要求が、デプロイメントコストをさらに削減するための主要なボトルネックとなっている。
この課題に対処するために、DeepSeek-V4.1-Flashという、552Bのバックボーンパラメータを持つマルチモーダルMixture-of-Experts(MoE)モデルを導入し、最大100万のトークンのコンテキストをサポートする。
Causal Encoder-Decoder (CED) アーキテクチャでは、デコード時にトークン毎の16Bパラメータを活性化するが、プリフィル時には8Bパラメータのみを活性化し、エージェントワークロードのコスト効率を大幅に向上させる。
KVキャッシュ圧縮の限界を押し上げるため、DeepSeek-V4.1-FlashはCompressed Sparse Attention 2 (CSA2)における層間KVキャッシュの再利用とFP4 KVキャッシュを組み合わせた。
これらの設計により、グローバルなKVキャッシュのフットプリント(通常はHBM)は、DeepSeek-V4-Flashのフットプリントのおよそ1/4のトークン当たり890バイトに削減される。
さらに、SWA境界リプレイと呼ばれる専用のデプロイメント最適化により、DeepSeek-V4.1-Flashは永続的なKVキャッシュのフットプリントを(SSDやホストメモリでは常に)DeepSeek-V4-Flashの約1/8に削減した。
KVキャッシュのフットプリントははるかに小さいが、ベースラインよりも大幅にパフォーマンスが向上している。
さらに、DeepSeek-V4アーキテクチャを合理化し、効率的なアーキテクチャ拡張をいくつか導入する。
我々は45Tトークンからなるマルチモーダルコーパス上でDeepSeek-V4.1-Flashを事前訓練し、総合的なポストトレーニングを行い、多様なテキストベースおよびマルチモーダルエージェントシナリオに強いパフォーマンスをもたらす。
モデルチェックポイントはhttps://huggingface.co/deepseek-ai/deepSeek-V4.1-Flashで入手できる。
関連論文リスト
- FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention [77.12062766962815]
Lookahead Sparse Attention (LSA)は、DeepSeek-V4アーキテクチャ上に構築されたNeural Memory Indexerを利用している。
このアーキテクチャをバックボーンフリーの非結合なトレーニング戦略でインスタンス化する。
FM-DS-V4は、物理KVキャッシュのフットプリントを、フルコンテキストベースラインのわずか13.5%まで圧縮することを示した。
論文 参考訳(メタデータ) (2026-06-08T06:25:54Z) - DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence [240.40422202055993]
We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language model。
DeepSeek-V4シリーズには、アーキテクチャと最適化のいくつかの重要なアップグレードが含まれている。
DeepSeek-V4シリーズは、長いコンテキストシナリオにおいて非常に効率的である。
論文 参考訳(メタデータ) (2026-04-26T14:49:33Z) - MiniCache: KV Cache Compression in Depth Dimension for Large Language Models [48.03117580340151]
キーバリュー(KV)キャッシュは、以前に生成されたトークンのキー値状態を格納する。
KVキャッシュのサイズはシーケンス長とともに線形に増加し、長いコンテキスト入力と広範囲なシーケンス生成を必要とするアプリケーションの課題を提起する。
レイヤ間のKVキャッシュを,新しい奥行きの観点から圧縮する,MiniCacheという,シンプルで効果的なアプローチを提案する。
論文 参考訳(メタデータ) (2024-05-23T09:43:52Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。