論文の概要: Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance
- arxiv url: http://arxiv.org/abs/2608.26480v1
- Date: Thu, 27 Aug 2026 00:11:41 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-28 16:30:58.204594
- Title: Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance
- Title(参考訳): LLM符号化性能向上のためのLedger-based Controlによるゼロショット自己組織化
- Abstract要約: 共有作業空間に対する管理作業者の足場の影響について検討する。
監督のオプス5は1回のパスで91%の得点を記録した。
マネジャーの実行はトークンの請求書を3倍にしますが、より大きなモデルに移行するよりも、より安価で精度が得られます。
- 参考スコア(独自算出の注目度): 2.8543100656957883
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Multi-agent large language model systems are widely reported to beat single-model baselines, but the evidence is mixed, and comparisons are usually confounded: pipelines change token budgets, tool calls, and prompts simultaneously, so an aggregate gain rarely reveals what actually helped. We investigate the effect of introducing the manager-worker scaffold over a shared filesystem workspace, with no training and no per-benchmark tuning, measured against the same model answering in a single pass. Across nine models -- five open-weight, spanning 9B to ~2.8T parameters, and four frontier closed models -- on the 100 latest hard LiveCodeBench problems, the scaffold's benefit is real but conditional: large and statistically significant for some (Qwen3.8-27B +23.4, GPT-5.6-Luna +10.6 and GPT-5.6-Terra +8.0, each over five paired passes; Kimi-K3 +30.4 and Minimax-M3 +11.0 over five paired passes with reasoning off, both at $p < 10^{-4}$, and +42 and +12 in a single pass at a 128k cap) and null or negative for others (Qwen3.6-35B -1 to -9 with reasoning off). With the manager, Opus-5 achieves the highest score in the study at 91% in one pass. Running a manager roughly triples the token bill, but it buys accuracy more cheaply than moving to a larger model does: GPT-5.6-Terra with a manager nearly matches Fable 5's single-call accuracy (85.0 against 87.4, $p = 0.59$) at a fifth of the price (\$11.71 against \$61.11 per 100-problem pass, $p < 10^{-4}$), and the Qwen-27B arm does it for \$51.75 on weights anyone can self-host. Our transcript analysis finds several mechanisms behind the gains, of which two recur: context management, in which short worker calls and shared notes organize state and reduce truncation, and problem decomposition. Improvements are modest for large models with reasoning enabled, but larger for some models with reasoning disabled and for smaller models with reasoning enabled.
- Abstract(参考訳): マルチエージェントの大規模言語モデルシステムは、シングルモデルベースラインを上回ると広く報告されているが、エビデンスが混ざり合っており、一般的には、パイプラインはトークンの予算を変更し、ツールコールとプロンプトを同時に変更する。
共有ファイルシステムのワークスペース上にマネージャと作業者の足場を導入し,トレーニングもベンチマーク単位のチューニングも行わず,同一のモデルに対して1回のパスで応答する。
9つのモデル – 9Bから2.8Tパラメータにまたがる5つのオープンウェイトと、100の最新のハードなLiveCodeBench問題において、フロンティアのクローズドモデル -- において、足場の利点は本物だが条件付きである: いくつか(Qwen3.8-27B +23.4、GPT-5.6-Luna +10.6、GPT-5.6-Terra +8.0、それぞれ5つのペアパス、Kimi-K3 +30.4、Minimax-M3 +11.0、および5つのペアパスは、それぞれ$p < 10^{-4}$と128kのパスで、+42と+12は負または負(Qwen3-35-B-19)である。
監督のオプス5は1回のパスで91%の得点を記録した。
GPT-5.6-Terra with a Manager with a Manager's single-call accuracy (85.0 against 87.4, $p = 0.59$) at a fifth of the price (\$11.71 against \61.11 per 100-problem pass, $p < 10^{-4}$) and the Qwen-27B arm do it for $51.75 on everyone can self-hosts。
我々のトランスクリプト分析は、この利得の背後にあるいくつかのメカニズムを見つけ、そのうちの2つの再帰的なコンテキスト管理は、短いワーカー呼び出しと共有ノートが状態を整理し、トランケーションを減らし、問題を分解するものである。
推論が有効になった大型モデルでは改善は控えめだが、推論が不可能なモデルではより大きく、推論が有効になったモデルではより小さいモデルではより大きい。
関連論文リスト
- Data Turnstile: A Scalable Open Framework for Function-Calling Data Generation [4.212647848163336]
本稿では,関数呼び出しのための高品質な合成トレーニングデータを生成する,オープンソースのフレームワークであるData Turnstileを紹介する。
Turnstileは、マルチターンツール使用のインタラクションを、バリデーションとエラーフィードバックループを備えた制約付き、ステップワイズ生成に分解する。
本論文では,Turnstileデータを用いた2つの関数呼び出しベンチマークにおけるドメイン適応の有効性を示す。
論文 参考訳(メタデータ) (2026-07-31T10:21:36Z) - Scaffolding the Strategist: Architecture-Dependent Reasoning Interventions in Hotelling Spatial Markets [0.01269104766024433]
構造化推論の介入が大規模言語モデルの戦略的経済的推論を改善するかどうかを検討する。
GPT-4.1-mini(標準命令追従モデル)とGPT-5-mini(推論最適化モデル)を5条件で評価した。
足場型とモデルアーキテクチャ間の統計的に有意な相互作用を見出した。
論文 参考訳(メタデータ) (2026-07-03T10:23:31Z) - ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents [61.184954205350984]
ClawArena-Teamは、41のマルチターン、マルチモーダル、マルチエージェントシナリオのベンチマークである。
全体的なスコア -- Subagent-Management Score (SMS) -- は、タスクの正しさを最小限のプライマリとモダリティのルーティング因子によって乗算する。
論文 参考訳(メタデータ) (2026-06-30T06:06:08Z) - When Do LLM Agents Treat Surface Noise Differently from Semantic Noise? A 68-Cell Measurement Study with a Held-Out Trace-Level Validation [9.055086193088083]
10大言語モデルによって駆動されるチェーン・オブ・シンクとReActエージェントに経験的現象を記述した。
平均的な摂動は、同等の厳しさのプレゼンテーション摂動よりも、最終的な答えを頻繁に変更する。
論文 参考訳(メタデータ) (2026-05-25T15:57:11Z) - What Do EEG Foundation Models Capture from Human Brain Signals? [64.48249643001402]
現代の脳波基礎モデルは、自己教師付き事前訓練を通じて生信号から直接学習する。
我々は3つのサブクエストに分解する: モデルが何を学習するか、モデルを何に使用するのか、そしてどのように説明できるのか。
3つの基礎モデル(CSBrain, CBraMod, LaBraM),5つの臨床タスク(MDD, Stress, ISRUC-Sleep, TUSL, Siena)と6ファミリー63機能レキシコンを含む。
論文 参考訳(メタデータ) (2026-05-12T01:57:53Z) - Attributing Emergence in Million-Agent Systems [68.53670424791751]
大規模言語モデル(LLM)は、個々のエージェントにおける人間のような推論と意思決定をシミュレートすることができる。
このような研究は、個々のエージェントにマクロな出現をもたらす必要がある。
Aumann--Shapley path-integral attribution to LLM-powered MAS at million-agent scale。
論文 参考訳(メタデータ) (2026-05-12T01:49:41Z) - ZenBrain: A Neuroscience-Inspired 7-Layer Memory Architecture for Autonomous AI Systems [51.56484100374058]
LongMemEval-500では、ZenBrainは長いコンテキストのオラクルのバイナリ・ジャッジの精度を4.5pp以内と一致させる。
ZenBrainは7層の神経科学にインスパイアされたメモリアーキテクチャである。
論文 参考訳(メタデータ) (2026-04-26T20:39:19Z) - The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs [9.989306175511238]
本稿では,RLHF+構成AIモデルと,より軽量に整合したモデルとを対比したケーススタディを提案する。
このギャップをアライメントフロアとして定義する: $_textfloor(m)=max_pS(m,p)-min_pS(m,p)$。
デプロイ時の監査指標として$_textfloor$を提案し、ペルソナのカスタマイズをデプロイする前に、小さなペルソナパネルで測定する。
論文 参考訳(メタデータ) (2026-04-10T08:04:41Z) - Humains-Junior: A 3.8B Language Model Achieving GPT-4o-Level Factual Accuracy by Directed Exoskeleton Reasoning [0.0]
Humans-Juniorは3.8Bモデルで、FACTS GroundingのパブリックサブセットのGPT-4oと$pm 5$ ppで一致している。
我々のアプローチは、最小指向の"Exoskeleton Reasoning"足場と、プロトコルコンプライアンスを教える振る舞いの微調整を組み合わせたものです。
論文 参考訳(メタデータ) (2025-10-29T20:12:36Z) - ProofAug: Efficient Neural Theorem Proving via Fine-grained Proof Structure Analysis [50.020850767257095]
本稿では,LLMに様々な粒度で自動化手法を付加するProofAugを提案する。
本手法は,オープンソースのDeep-math-7bベースモデルとIsabelle証明アシスタントを用いて,MiniF2Fベンチマークで検証した。
また、ProofAugのLean 4バージョンを実装し、Kimina-Prover-seek-Distill-1.5Bのパス@1のパフォーマンスを44.3%から50.4%に改善します。
論文 参考訳(メタデータ) (2025-01-30T12:37:06Z) - Improving Large Models with Small models: Lower Costs and Better Performance [81.55672406002715]
我々は,小型モデルと大規模モデルの協調のための一般的なパラダイムであるData Shunt$+$ (DS$+$)を提案する。
例えば、ChatGPTはAmazon Productの感情分析で9,43%の精度を達成し、DS$+は9,5.64%の精度を達成している。
論文 参考訳(メタデータ) (2024-06-15T14:44:43Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。