論文の概要: OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling
- arxiv url: http://arxiv.org/abs/2608.05141v1
- Date: Wed, 05 Aug 2026 17:58:15 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-06 14:48:44.073134
- Title: OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling
- Title(参考訳): OctoLong: クロスリポジトリなコードコンテキストが長期のモデリングを促進する
- Authors: Indraneil Paul, Falko Helm, Goran Glavaš, Iryna Gurevych,
- Abstract要約: 既存の長文コーパスは書籍、学術論文、コードリポジトリによって支配されている。
AST、言語サーババックエンド、パッケージマネージャを備えたコンテキストエンジニアリングパイプラインであるOctoLongを紹介します。
我々は600Mから14Bのパラメータのベースモデルから派生した、有能な長文オープンなLMのスイートであるOctoLong-Instructを訓練する。
従来の文脈拡張コーパスの12%をOctoLongデータで置き換えただけで、かなりの利益が得られます。
- 参考スコア(独自算出の注目度): 48.07406452417886
- License: http://creativecommons.org/licenses/by-sa/4.0/
- Abstract: Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, self-improvement, and long-horizon agentic workflows. Existing long-context corpora, however, are dominated by books, academic articles, and code repositories, which are finite resources and often scarce in long-distance dependencies. In this work, we introduce OctoLong, a context engineering pipeline that instruments an AST parser, a language server backend, and a package manager to facilitate the recursive retrieval of code references, enabling the curation of dependency-rich code contexts of millions of tokens in length. We then train OctoLong-Instruct, a suite of capable long-context open LMs, derived from base models ranging in size from 600M to 14B parameters, via context-extension mid-training on a ~50B-token mixture containing ~6.2B tokens of OctoLong code contexts, followed by ~10B tokens of instruction tuning. Our training ablations and experimental evaluations against 18 state-of-the-art open-weight long-context LMs show that supplanting just 12% of traditional context-extension corpora with OctoLong data yields substantial gains in long-range retrieval, long-term state tracking, repository-level code understanding, and downstream agentic tasks, while also enhancing API usage in short-context coding scenarios.
- Abstract(参考訳): 言語モデル(LM)の文脈長は、文脈内学習、自己改善、長期エージェントワークフローの要求によって劇的に増加してきた。
しかし、既存の長いコンテキストのコーパスは書籍、学術論文、コードリポジトリに支配されている。
本研究では,ASTパーサ,言語サーババックエンド,パッケージマネージャを備えたコンテクストエンジニアリングパイプラインであるOctoLongを導入し,コード参照の再帰的検索を可能にし,数百万のトークンの依存性に富んだコードコンテキストのキュレーションを可能にする。
次にOctoLong-Instructをトレーニングし、OctoLongのコードコンテキストの6.2Bトークンを含む約50B-token混合上で、600Mから14Bパラメータのベースモデルから派生した、有能な長文オープンLMのスイートであるOctoLong-Instructをトレーニングする。
従来の文脈拡張コーパスの12%をOctoLongデータに置き換えることで、長期検索、長期状態追跡、リポジトリレベルのコード理解、下流のエージェントタスクにおいて大幅に向上すると同時に、短文コーディングシナリオにおけるAPI使用率の向上も実現している。
関連論文リスト
- Bootstrap Your Own Context Length [74.61148597039248]
長文言語モデルを学習するためのブートストラップ手法を提案する。
提案したデータ合成ワークフローは、短いコンテキスト言語モデル、テキスト検索、文書収集のみを必要とする。
我々は,オープンソースのLlama-3ファミリを用いて実験を行い,最大100万トークンまでコンテキスト長を拡張できることを実証した。
論文 参考訳(メタデータ) (2024-12-25T10:08:54Z) - How to Train Long-Context Language Models (Effectively) [75.5418485597276]
言語モデル(LM)の継続学習と教師付き微調整(SFT)を行い,長文情報の有効利用について検討した。
コードリポジトリと書籍は長いデータの優れた情報源であることがわかったが、それらと高品質の短文データを組み合わせることが不可欠である。
最終モデルであるProLong-8Bは、128Kの同様のサイズのモデル間で、最先端の長文性能を示す。
論文 参考訳(メタデータ) (2024-10-03T16:46:52Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。