論文の概要: SepSeq: A Training-Free Framework for Long Numerical Sequence Processing in LLMs
- arxiv url: http://arxiv.org/abs/2604.07737v1
- Date: Thu, 09 Apr 2026 02:40:34 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-04-10 18:34:05.647768
- Title: SepSeq: A Training-Free Framework for Long Numerical Sequence Processing in LLMs
- Title(参考訳): SepSeq: LLMにおける長期数値シーケンス処理のためのトレーニングフリーフレームワーク
- Authors: Jie Sun, Yu Liu, Lu Han, Qiwen Deng, Xiang Shu, Yang Xiao, Xingyu Lu, Jun Zhou, Pengfei Liu, Lintao Ma, Jiancan Wu, Xiang Wang,
- Abstract要約: 本稿では,セパレータトークンを戦略的に挿入することで分散を緩和する学習自由なプラグアンドプレイフレームワークを提案する。
メカニカルには、セパレータトークンが注目シンクとして機能し、グローバルなコンテキストを維持しながら、局所的なセグメントに注意を向けることが示される。
- 参考スコア(独自算出の注目度): 51.84813231879529
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: While transformer-based Large Language Models (LLMs) theoretically support massive context windows, they suffer from severe performance degradation when processing long numerical sequences. We attribute this failure to the attention dispersion in the Softmax mechanism, which prevents the model from concentrating attention. To overcome this, we propose Separate Sequence (SepSeq), a training-free, plug-and-play framework to mitigate dispersion by strategically inserting separator tokens. Mechanistically, we demonstrate that separator tokens act as an attention sink, recalibrating attention to focus on local segments while preserving global context. Extensive evaluations on 9 widely-adopted LLMs confirm the effectiveness of our approach: SepSeq yields an average relative accuracy improvement of 35.6% across diverse domains while reducing total inference token consumption by 16.4% on average.
- Abstract(参考訳): 変圧器をベースとした大規模言語モデル(LLM)は理論的には大量のコンテキストウィンドウをサポートするが、長い数値列を処理する際には厳しい性能劣化に悩まされる。
我々は、この故障をSoftmaxメカニズムの注意分散に起因し、モデルが注意を集中することを防ぐ。
これを解決するために,セパレータトークンを戦略的に挿入することで分散を緩和する,トレーニング不要なプラグアンドプレイフレームワークであるセパレートシーケンス(SepSeq)を提案する。
メカニカルには、セパレータトークンが注目シンクとして機能し、グローバルなコンテキストを維持しながら、局所的なセグメントに注意を向けることが示される。
SepSeqは、様々な領域で平均35.6%の相対的精度向上を達成し、総推論トークン消費量を平均16.4%削減します。
関連論文リスト
- SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator [65.62084602011596]
大規模言語モデル(LLM)は、自然言語処理タスクの範囲で例外的な性能を示した。
特定の意味のないセパレータトークン(句読点)は意味的に意味のあるトークンと比較して注意点に不均等に寄与する。
SepLLMは,これらのセグメントを圧縮し,冗長なトークンを除去することによって推論を高速化する,プラグアンドプレイフレームワークである。
論文 参考訳(メタデータ) (2024-12-16T18:58:57Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。