論文の概要: Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale
- arxiv url: http://arxiv.org/abs/2606.15079v1
- Date: Sat, 13 Jun 2026 03:21:49 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-06-16 16:21:32.769803
- Title: Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale
- Title(参考訳): Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale
- Abstract要約: 我々はLing-2.6とRing-2.6を紹介した。
Ling-2.6は、即時応答生成と出力トークン当たりの高機能に最適化されている。
Ring-2.6はより深い推論とより高度なエージェント用に調整されている。
- 参考スコア(独自算出の注目度): 231.1015617648021
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: Efficient and scalable agentic intelligence requires models that can deliver both low-latency responses and strong reasoning capabilities while remaining practical to train, serve, and deploy. In this report, we present Ling-2.6 and Ring-2.6, a family of models designed to address this challenge at scale. Ling-2.6 is optimized for instant response generation and high capability per output token, whereas Ring-2.6 is tailored for deeper reasoning and more advanced agentic workflows. Instead of training from scratch, we upgrade the Ling-2.0 base model through architectural migration pre-training and large-scale post-training. This upgrade is guided by a unified co-design of model architecture, optimization objectives, serving systems, and agent training environments, enabling improvements in both model capability and deployment efficiency. At the architectural level, we introduce a hybrid linear attention design that integrates Lightning Attention with MLA, improving the efficiency of long-context training and decoding. To further enhance token efficiency, we optimize capability per output token through Evolutionary Chain-of-Thought, Linguistic Unit Policy Optimization, bidirectional preference alignment, and shortest-correct-response distillation. For agentic capabilities, we propose KPop, a reinforcement learning framework designed to support stable training of Ring-2.6-1T on large-scale environment-grounded data. KPop improves training efficiency through asynchronous scheduling across coding, search, tool use, and workflow execution, enabling scalable learning from complex agent-environment interactions. Together, Ling-2.6 and Ring-2.6 provide a practical pathway toward efficient, scalable, and open agentic systems. We open-source all checkpoints in the 2.6 family to support further research and development in practical agentic intelligence.
- Abstract(参考訳): 効率的でスケーラブルなエージェントインテリジェンスには、トレーニング、サービス、デプロイに実用的でありながら、低レイテンシ応答と強力な推論能力の両方を提供するモデルが必要です。
本稿では,Ling-2.6とRing-2.6について述べる。
Ling-2.6は即時応答生成と出力トークン毎の高機能に最適化されているのに対し、Ring-2.6はより深い推論とより高度なエージェントワークフローに最適化されている。
ゼロからトレーニングする代わりに、アーキテクチャ移行前トレーニングと大規模ポストトレーニングを通じてLing-2.0ベースモデルをアップグレードします。
このアップグレードは、モデルアーキテクチャ、最適化目標、サービスシステム、エージェントトレーニング環境の統一された共同設計によって導かれる。
アーキテクチャレベルでは、Lightning Attention と MLA を統合したハイブリッド線形アテンション設計を導入し、長期学習と復号化の効率を向上させる。
トークン効率をさらに高めるため,進化的連鎖・言語単位政策最適化,双方向優先調整,最短誤り応答蒸留を通じて,出力トークン当たりの性能を最適化する。
KPopは大規模環境下でのRing-2.6-1Tの安定トレーニングを支援するための強化学習フレームワークである。
KPopは、コーディング、検索、ツール使用、ワークフロー実行といった非同期スケジューリングを通じて、トレーニング効率を改善し、複雑なエージェントと環境のインタラクションからスケーラブルな学習を可能にする。
Ling-2.6とRing-2.6は、効率的でスケーラブルでオープンなエージェントシステムへの実践的な経路を提供する。
我々は2.6ファミリーのすべてのチェックポイントをオープンソース化し、実用的なエージェントインテリジェンスの研究と開発を支援します。
関連論文リスト
- ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search [54.43696074323815]
ZGCM-1は、極端なデータ、システム、アルゴリズムの効率でゼロから訓練された、完全にオープンな7Bファンデーションモデルである。
エージェントがクラスタ操作を自律的に管理する、AIネイティブなR&Dワークフローを確立します。
ZGCM-1-7Bは一般的なベンチマークで7Bモデルファミリ間で競合する。
論文 参考訳(メタデータ) (2026-09-11T17:18:04Z) - MARS: Optimizing Dual-System Deep Research via Multi-Agent Reinforcement Learning [82.14973479594367]
複雑な推論タスクのための大規模言語モデル(LLM)は、直感的で意図的な認知プロセスを橋渡しする革新的なアプローチを必要とする。
本稿では,Multi-Agent System for Deep ReSearch (MARS)を提案する。
論文 参考訳(メタデータ) (2025-10-06T15:42:55Z) - Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation [65.3648667980258]
視覚言語モデル(VLM)に基づくGUIエージェントは複雑なタスクの自動化を約束するが、強化学習(RL)の適用において大きな課題に直面している。
異種モジュールを高度に非結合的に協調するGUIエージェントのための非結合エージェントRLトレーニングフレームワークであるDARTを提案する。
OSWorldのベンチマークでは、DART-GUI-7Bは42.13%のタスク成功率、14.61%の絶対ゲイン、オープンソースSOTAよりも7.34%高い。
論文 参考訳(メタデータ) (2025-09-28T13:19:20Z) - Revolutionizing Reinforcement Learning Framework for Diffusion Large Language Models [49.911784762244814]
TraceRLは拡散言語モデル(DLM)のための軌道対応強化学習フレームワークである
我々は最先端の拡散言語モデル、すなわち TraDo を導出する。
TraDo-8B-InstructはQwen2.5-7B-Instructで6.1%、Llama3.1-8B-Instructで51.3%の精度向上を実現している。
論文 参考訳(メタデータ) (2025-09-08T17:58:06Z) - Oranits: Mission Assignment and Task Offloading in Open RAN-based ITS using Metaheuristic and Deep Reinforcement Learning [10.571108374756184]
オープンLAN(Open RAN)ベースのインテリジェントトランスポートシステム(ITS)におけるミッション割り当てとタスクオフロードについて検討する。
既存の研究はしばしば、ミッション間の複雑な相互依存と、エッジサーバへのタスクのオフロードに伴うコストを見落としている。
我々は、オーラニッツ(Oranits)という、ミッション依存とオフロードコストを明示的に考慮し、車両の協調によって性能を最適化する新しいシステムモデルを紹介した。
論文 参考訳(メタデータ) (2025-07-25T23:13:09Z) - Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs [51.21041884010009]
Ring-liteは、強化学習(RL)により最適化されたMixture-of-Experts(MoE)ベースの大規模言語モデルである
我々のアプローチは、挑戦的なベンチマーク上でのSOTA(State-of-the-art)の小規模推論モデルの性能と一致する。
論文 参考訳(メタデータ) (2025-06-17T17:12:34Z) - BiERL: A Meta Evolutionary Reinforcement Learning Framework via Bilevel
Optimization [34.24884427152513]
双レベル最適化(BiERL)による一般的なメタERLフレームワークを提案する。
我々は、内部レベルの進化した経験を情報的人口表現に組み込むエレガントなメタレベルアーキテクチャを設計する。
我々は MuJoCo と Box2D タスクの広範な実験を行い、一般的なフレームワークとして BiERL が様々なベースラインを上回り、ERL アルゴリズムの多様性の学習性能を一貫して向上することを検証する。
論文 参考訳(メタデータ) (2023-08-01T09:31:51Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。