論文の概要: Frontier Autolab: Organizational Memory, Adversarial Dissent and Temporal Leakage in Multi-Agent LLM Firms Across Fifty Years of Technological Change
- arxiv url: http://arxiv.org/abs/2609.36739v1
- Date: Tue, 29 Sep 2026 05:12:12 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-30 21:28:47.204123
- Title: Frontier Autolab: Organizational Memory, Adversarial Dissent and Temporal Leakage in Multi-Agent LLM Firms Across Fifty Years of Technological Change
- Title(参考訳): Frontier Autolab:50年にわたる技術変化の多分野LPM企業における組織記憶, 対人距離, 時間漏洩
- Abstract要約: フロンティア・オートラブ(Frontier Autolab)は、1990年から2040年までの9つの技術時代において、シミュレーションされた会社を再び設立しなければならない、長期にわたるテストベッドである。
6つの時代は歴史に対して、1つはライブ市場に対して、2つは公開予測である。
歴史的に評価された24のすべての時代において、判事は、同社がどの場所を建設するかという選択よりも、来るべきシフトを認識していると評価した。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Multi-agent LLM systems are increasingly structured like organizations, with roles, critics and shared memory, yet they are evaluated on tasks that last minutes. We ask how such an organization behaves when the ground it stands on keeps moving. Frontier Autolab is a long-horizon testbed in which one simulated firm, voiced by sixteen role personas and a dedicated Red Team, must re-found itself in nine technology eras from 1990 to 2040. Each era is temporally gated: the firm decides from a dated briefing, a historian-judge then reveals what happened and scores the decision on a five-dimension rubric, and lessons enter a persistent Playbook. Six eras are scored against history, one against the live market and two are open forecasts. Across four trajectories (36 era decisions, 180 subscores) we find a consistent foresight-commitment gap: in all 24 historically scored eras the judge rated the firm's recognition of the coming shift above its choice of where to build (mean gap 1.9 points on a 10-point scale), because boards chose the layer their existing assets could reach. Organizational design shaped long-run character. A Red Team armed with numeric kill gates produced fifty years of gated pilots and no product, and the rubric rated this firm highest; firms whose memory stored market-structure lessons pivoted every era, while a firm whose memory stored only validation procedure kept one method throughout. We also show why such results are hard to trust. Scores rise across eras in every run while the judge's own hindsight subscore falls (within-run r = -0.58), so apparent learning is confounded with recall of history, and we trace further distortions to self-judging, briefing selection and score aggregation. We release all records and an API harness, and specify fictional and post-cutoff eras that would turn the testbed into a benchmark.
- Abstract(参考訳): マルチエージェントLLMシステムは、役割、批評家、共有メモリを持つ組織のようにますます構造化されているが、最後の数分で評価される。
そのような組織は、地上が動き続けるとき、どのように振る舞うかを尋ねます。
フロンティア・オートラブ(Frontier Autolab)は、1990年から2040年までの9つのテクノロジー時代において、16人のロール・ペルソナと1人のレッド・チームによって声を掛けられたシミュレーション会社である。
会社は、日付のついたブリーフィングから決定し、歴史家の判断により、何が起こったかを明らかにし、5次元のルーリックで決定を下し、レッスンが永続的なプレイブックに入る。
6つの時代は歴史に対して、1つはライブ市場に対して、2つは公開予測である。
4つの軌跡(36年代決定、180サブスコア)にわたって、一貫した監視・制約のギャップを見いだす: 歴史的に評価されたすべての24つの時代において、裁判官は、同社がどの場所に建設するかという選択よりも、次のシフトを認識できると評価した(平均10ポイントの差1.9ポイント)。
組織デザインのロングランキャラクタ。
数字のキルゲートで武装したレッドチームが50年分のゲートパイロットを生産し、製品も生産しなかった。
また、なぜそのような結果が信用し難いのかも示します。
審査員自身の後見サブスコアが低下する(r = -0.58で)間、各実行中の時代にわたってスコアが上昇するので、明らかな学習は歴史の記憶と組み合わせられ、自己判断、ブリーフィング選択、スコアアグリゲーションにさらなる歪みを辿る。
すべてのレコードとAPIハーネスをリリースし、テストベッドをベンチマークにするフィクションとポストカットの期間を指定します。
関連論文リスト
- TWIST: A Proposed Benchmark for Intervention Quality in Conversational Memory, with a Human-Validated Draft-Alignment [0.0]
TWISTは、補完的かつ未測定なプロパティ、介入品質のためのベンチマークスイートである。
4つのトラックは、未解決の緊張検知をカバーし、現在の信念に応えて、記録に反するドラフトを検証している。
システムTWISTプロファイルはリコールスコアの他に、メモリがいつ介入すべきかを計測する。
論文 参考訳(メタデータ) (2026-09-23T12:25:28Z) - CART: Closed-Loop Adaptive Red Teaming for Large Language Models [91.66397227953821]
CART(Closed-Loop Adaptive Red Teaming)は、各結果を使用して、次にテストするものをガイドするフレームワークである。
テストを生成するChallengeと、テスト中のTarget、テキストのみのモデルや境界ツール使用エージェント、結果を評価するJuiceを分離する。
論文 参考訳(メタデータ) (2026-09-23T04:16:46Z) - Agent Zero Memory: Provenance-Aware Long-Term Memory for LLM Agents [3.6666917625778193]
Agent Zero Memory(エージェントゼロメモリ)は、長期記憶システムである。
ユーザの会話、ファイル、接続されたソースを3つの並列メモリシステムに消去する。
論文 参考訳(メタデータ) (2026-08-30T06:55:59Z) - Evaluation Choices Decide the Forecasting Leaderboard: Evidence from a Production Marketplace Panel [0.0]
我々は,2025年代の時系列基礎モデルを含む24の予測手法と1つの教科書参照をベンチマークした。
我々は、データ、地平線、期間を固定し、評価設計だけを変える。
3つの選択はそれぞれ、見出しの結論を逆転または解消する。
論文 参考訳(メタデータ) (2026-08-20T12:38:39Z) - Quipu: A Governed Bitemporal Knowledge Graph Store [0.0]
Quipuは4つのデフォルトを逆転する組み込み可能なナレッジストアだ。
議事録が保留中のポストステートを評価するゲートを経由する以外は、事実は入らない。
支配的な作家からの痕跡は、その修復によって監査名のライブ執行のギャップを表面化する。
論文 参考訳(メタデータ) (2026-08-17T17:04:29Z) - OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models [87.08963239083614]
コンピュータ利用エージェント(CUA)はデジタル世界で急速に進歩している。
この分野は、CUA軌道の審査員として、視覚言語モデル(VLM)に変わりつつある。
我々は,CUA軌道上のVLM判定値を評価する,現実的で高品質なベンチマークOSRewardを紹介する。
論文 参考訳(メタデータ) (2026-07-30T17:57:41Z) - Less Context, More Accuracy: A Bi-Temporal Memory Engine for LLM Agents Where a Lean Retrieved Context Beats the Full History [2.5401434059780468]
Engramは、バイテンポラルデータモデル上のデュアルプロセスメモリエンジンである。
高速書き込みパスは、クリティカルパスにLSMなしでエピソードを付加する。
ハイブリッドリードパスは、密度、語彙、グラフ、および電流/セイレンス信号を融合する。
論文 参考訳(メタデータ) (2026-06-05T11:43:56Z) - WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction [72.1620416874118]
マルチモーダルな言語モデルは、長距離エージェントとしてますます多くデプロイされている。
既存のベンチマークは、静的対話上のリコールを測定し、メモリを1つのタスクの精度に分解し、キャプションに対する視覚的な観察を減らす。
マルチモーダルエージェントメモリを,観測可能な4段階ライフサイクルを持つアクションワールドインタラクションループとして定式化する。
論文 参考訳(メタデータ) (2026-05-28T04:27:20Z) - RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies [54.23445842621374]
記憶は、長い水平と歴史に依存したロボット操作にとって重要である。
近年,視覚言語アクション(VLA)モデルにメモリ機構が組み込まれ始めている。
本稿では,VLAモデルの評価と進展のための大規模標準ベンチマークであるRoboMMEを紹介する。
論文 参考訳(メタデータ) (2026-03-04T21:59:32Z) - The Pensieve Paradigm: Stateful Language Models Mastering Their Own Context [48.70817145536136]
StateLMは、自身の状態を管理するための内部推論ループを備えた、新しいファンデーションモデルのクラスである。
動的に自分自身のコンテキストを設計することを学ぶことで、私たちのモデルは固定された窓のアーキテクチャの監獄から解放されます。
論文 参考訳(メタデータ) (2026-02-12T16:00:01Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。