論文の概要: LongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction Synthesis
- arxiv url: http://arxiv.org/abs/2607.06160v1
- Date: Tue, 07 Jul 2026 11:35:48 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-08 21:24:51.503392
- Title: LongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction Synthesis
- Title(参考訳): LongCrafter:Evidence-Graph-Guided Instruction Synthesisによる多言語長文理解を目指して
- Authors: Chenhao Yuan, Yinhao Xu, Shuwen Xu, Xizhi Yang, Jiaxiang Liu, Chenxi Zhou, Shaoping Huang, Haolin Ren, Pengfei Cao, Jun Zhao, Kang Liu,
- Abstract要約: LongCrafterは階層的なタスク分類とエビデンスベースのパイプラインを結合する構造化合成フレームワークである。
LongCrafterはタスク整列した長いコンテキストを構築し、それらをパラグラフの依存関係をモデル化する明示的なエビデンスグラフに分解し、位置するエビデンスに厳格に根ざした命令応答ペアを生成する。
LongCrafterのデータに基づいて微調整されたモデルは、すべてのSFTベースラインや、公式の訓練後のモデルよりも優れています。
- 参考スコア(独自算出の注目度): 25.64387650835006
- License: http://creativecommons.org/licenses/by-sa/4.0/
- Abstract: Synthesizing long-context supervised fine-tuning (SFT) data is a scalable way to enhance the long-context understanding of large language models (LLMs), yet existing approaches share three limitations: narrow task coverage, insufficient instruction difficulty, and a lack of faithfulness supervision. We propose \textbf{LongCrafter}, a structured synthesis framework that couples a hierarchical task taxonomy with an evidence-grounded pipeline. The taxonomy organizes long-context understanding into local/shallow and global/deep levels and yields 32 fine-grained task types that serve as a global generative prior. Guided by this taxonomy, LongCrafter constructs task-aligned long contexts, decomposes them into explicit evidence graphs that model cross-paragraph dependencies, and generates instruction--response pairs strictly grounded in the located evidence spans, ensuring both controllable difficulty and faithful, traceable reasoning. Models fine-tuned on LongCrafter data outperform all SFT baselines and even the official post-trained models on LongBench, LongBench~v2, and LooGLE across both Qwen2.5-7B and LLaMA-3.1-8B, with the largest gains on high-difficulty tasks. Further analysis shows that LongCrafter data is more diverse and better spread across difficulty levels, and that the trained models locate evidence robustly regardless of position, effectively mitigating the ``lost in the middle'' problem.
- Abstract(参考訳): 長期コンテキスト教師付き微調整(SFT)データを合成することは、大規模言語モデル(LLM)の長期コンテキスト理解を強化するスケーラブルな方法であるが、既存のアプローチでは3つの制限がある。
本稿では,階層的なタスク分類とエビデンス基底パイプラインを結合した構造化合成フレームワークである‘textbf{LongCrafter} を提案する。
分類学は、局所・浅層・大域・深層への長文の理解を組織し、グローバルな生成先行として機能する32のきめ細かいタスクタイプを生成する。
この分類によって導かれたロングクラフトは、タスク整列した長いコンテキストを構築し、それらをクロスパラグラフの依存関係をモデル化する明示的なエビデンスグラフに分解し、位置するエビデンスに厳格に根ざした命令応答ペアを生成する。
ロングクラフトのデータに基づいて微調整されたモデルは、全てのSFTベースラインや、Qwen2.5-7BとLLaMA-3.1-8Bの両方にわたるLongBench、LongBench~v2、LooGLEのトレーニング後の公式モデルよりも優れており、高い微分タスクで最大の利益を得ている。
さらなる分析により、LongCrafterのデータはより多様性があり、難易度に分散し、トレーニングされたモデルは、位置に関係なくしっかりと証拠を見つけ、事実上「中間のロスト」問題を緩和することを示した。
関連論文リスト
- IS-CoT: Breaking the Long-form Generation Collapse via Interleaved Structural Thinking [50.381430711141995]
本稿では,Interleaved Structure Chain-of-Thought (IS-CoT) フレームワークを紹介する。
IS-CoTは動的Plan-Write-Reflectサイクルを生成プロセスに埋め込む。
我々は、多教師パイプラインを介して、インターリーブされた推論トレースの高品質なデータセットを構築し、IS-Writer-8Bを訓練する。
論文 参考訳(メタデータ) (2026-06-08T16:31:00Z) - DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding [63.257540233507626]
本稿では、構造化解析、局所化、推論のワークフローを実行するためにモデルを必要とするパラダイムを提案する。
ショートページトレーニングから超長文書への堅牢な一般化を示し、視覚的検索・拡張生成システムと自然に相乗効果を示す。
論文 参考訳(メタデータ) (2026-04-14T14:39:26Z) - WildLong: Synthesizing Realistic Long-Context Instruction Data at Scale [86.25450054683172]
WildLongは、実際のユーザクエリからメタ情報を取り出して、スケーラブルなデータを生成する。
クロスドキュメント比較やアグリゲーションといったマルチドキュメント推論をサポートする。
ベンチマーク全体で、既存のオープンソースの長期コンテキスト最適化モデルを上回っている。
論文 参考訳(メタデータ) (2025-02-23T18:59:09Z) - Generalizing From Short to Long: Effective Data Synthesis for Long-Context Instruction Tuning [103.65680870130839]
本研究では,長期事前学習モデルの学習後段階の指導データを設計する方法について検討する。
制御された研究では、短い文脈で調整されたモデルが、より長いコンテキストに効果的に一般化できることが判明した。
これらの知見に基づいて,新しいデータ合成フレームワークであるコンテキスト合成を提案する。
論文 参考訳(メタデータ) (2025-02-21T17:02:40Z) - LongFaith: Enhancing Long-Context Reasoning in LLMs with Faithful Synthetic Data [19.79929012055293]
LongFaithは忠実な長文推論命令データセットを合成するための新しいパイプラインである。
基礎的真理と引用に基づく推論のプロンプトを統合することにより、注意散らしを排除し、推論連鎖の精度を向上させる。
論文 参考訳(メタデータ) (2025-02-18T06:40:23Z) - GATEAU: Selecting Influential Samples for Long Context Alignment [59.579128690086385]
GATEAUは、長距離依存関係に富む影響力のあるサンプルを同定する。
選択されたサンプルに基づいて訓練されたモデルは、より良い指示追従と長文理解能力を示す。
論文 参考訳(メタデータ) (2024-10-21T04:30:53Z) - Long Context Alignment with Short Instructions and Synthesized Positions [56.1267385315404]
本稿では,ステップスキッピングアライメント(SkipAlign)を紹介する。
これは、Large Language Models(LLMs)の長期コンテキスト機能を強化するために設計された新しい技術である。
ベースモデルとアライメントデータセットを慎重に選択することで、SkipAlignは6Bパラメータだけで最高のパフォーマンスを実現し、LongBenchのGPT-3.5-Turbo-16Kのような強力なベースラインに匹敵する。
論文 参考訳(メタデータ) (2024-05-07T01:56:22Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。