論文の概要: Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
- arxiv url: http://arxiv.org/abs/2608.31075v2
- Date: Tue, 01 Sep 2026 03:17:42 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-02 16:31:35.702039
- Title: Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
- Title(参考訳): 人間のスーパービジョンを超えて大規模な推論モデルをスケールする:超知能への道
- Authors: Zhiqin Yang, Jingwen Fu, Yuhan Liu, Hengyu Liu, Yonggang Zhang, Kainan Cao, Zizhuo Zhang, Chenxin Li, Ruibin Yuan, Jiahao Pan, Jiankai Sun, Zhenyuan Zhang, Yibo Li, Yunlong Lin, Jing Xiong, Sida Lin, Bo Han, Wei Xue, Yike Guo,
- Abstract要約: 検証可能な報酬による強化学習は、数学やコードの推論を大幅に改善することができる。
しかし、人間の直接監督は、モデル生成体験のスケールと複雑さに追随することはできない。
本稿では,人間の指導が学習ループから徐々に遠ざかっていくにつれて,LRMがいかに改善していくかを検討する。
- 参考スコア(独自算出の注目度): 77.303673110294
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks remains difficult because reliable rewards are harder to obtain and direct human supervision cannot keep pace with the scale and complexity of model-generated experience. This paper studies how LRMs can continue to improve as human supervision gradually recedes from the learning loop. We examine two connected dimensions of this problem. The reward axis traces the development from per-instance human judgments to reusable verifiers and rewards that operate even without human feedback. The experience axis examines how learning can progress from human-curated tasks and environments toward self-generated curricula, constructed environments, and autonomous co-evolution. We connect these dimensions through a five-level ladder from L0 to L4 that identifies which parts of the learning process remain under continued human control. Our analysis further highlights the risks introduced by increasingly autonomous rewards and experience generation, including reward hacking, feedback drift, curriculum collapse, and environment errors. Consequently, we also provide the evaluation around three complementary objects: policy capability, feedback fidelity, and experience quality. This analysis provides a structured account of current approaches to scaling LRMs beyond human supervision and the open problems involved in developing self-sustaining learning systems toward superintelligence. Furthermore, we maintain a continuously updated \href{https://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision}{GitHub repository} to track the latest advances.
- Abstract(参考訳): 大規模推論モデル(LRM)の最近の進歩は、検証可能な報酬(RLVR)による強化学習が、数学やコードの推論を大幅に改善し、結果を自動的にチェックできることを示した。
信頼性の高い報酬を得るのが難しく、人間の直接監督は、モデル生成体験の規模や複雑さに追随できないため、この進歩をオープンエンドおよびエージェント的タスクに拡張することは依然として困難である。
本稿では,人間の指導が学習ループから徐々に遠ざかっていくにつれて,LRMがいかに改善していくかを検討する。
この問題の2つの連結次元について検討する。
報酬軸は、インスタンス当たりの人間の判断から、人間のフィードバックがなくても動作する再利用可能な検証や報酬へと発展を辿る。
体験軸は、人間の計算したタスクや環境から、自己生成したカリキュラム、構築された環境、そして自律的な共進化へと学習がどのように進むかを調べる。
我々はこれらの次元をL0からL4までの5段階のはしごを通して接続し、学習過程のどの部分が継続的な人間の制御下にあるかを特定する。
我々の分析は、報奨ハッキング、フィードバックドリフト、カリキュラム崩壊、環境エラーなど、ますます自律的な報酬や経験生成によってもたらされるリスクをさらに強調している。
その結果,政策能力,フィードバック忠実度,経験的品質の3つの相補的対象について評価を行った。
この分析は、人間の監督を超えてLEMをスケールするための現在のアプローチと、超知性に向けた自己維持学習システムの開発に関わるオープンな問題について、構造化された説明を提供する。
さらに、最新の進歩を追跡するために継続的に更新された \href{https://github.com/visitworld123/Awesome-Scaling-LRM-Beyond-Human-Supervision}{GitHub リポジトリも維持しています。
関連論文リスト
- Aspire: Can Models Self-Evolve from Vague Goals? [61.65007006747275]
我々はあいまいなゴール駆動型自己進化のベンチマークであるASPIREを紹介する。
ASPIREは、統一された対話環境におけるモデルウェイトとエージェントハーネスの進化をサポートする。
実験の結果,曖昧な目標が探索努力を目標解釈に向けることが示唆された。
論文 参考訳(メタデータ) (2026-08-31T17:14:59Z) - Training Skills Like Parameters via Self-Supervised Semantic Diffusion [15.133304131514713]
大規模言語モデル(LLM)は、創造的なスクリーンライティングなど、高度に専門化されたオープンエンドドメインの専門家には不足することが多い。
本稿では,拡散モデルの破壊・再構成パラダイムに着想を得た,新規で教師なしの自己進化型エージェントフレームワークを提案する。
ショート・ドラマ・スクリーンライティングの課題について,我々の枠組みを評価した。
論文 参考訳(メタデータ) (2026-07-30T01:03:59Z) - Guided Self-Evolving LLMs with Minimal Human Supervision [53.111086364268566]
無誘導の自己進化システムは、しばしば訓練として素早く、または劣化する。
R-Fewはガイド付きセルフプレイチャレンジャー(Self-Play Challenger)買収フレームワークで、コンテキスト内接地と混合トレーニングを通じて、軽量な人間の監視を取り入れている。
R-Fewは、数学と一般的な推論ベンチマークで一貫した反復的な改善を実現している。
論文 参考訳(メタデータ) (2025-12-02T07:06:11Z) - Absolute Zero: Reinforced Self-play Reasoning with Zero Data [57.30662797376754]
検証可能な報奨付き強化学習(RLVR)は,大規模言語モデルの推論能力を高めることを約束している。
本稿では,AZR(Absolute Zero Reasoner)について紹介する。
AZRは、コーディングおよび数学的推論タスクにおける全体的なSOTA性能を達成し、既存のゼロセットモデルより優れている。
論文 参考訳(メタデータ) (2025-05-06T09:08:00Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。