論文の概要: DreamLedger: Where to Refuse World-Model Imagination Using Execution-Settled Credit
- arxiv url: http://arxiv.org/abs/2608.23863v2
- Date: Wed, 26 Aug 2026 20:22:12 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-28 14:16:51.759138
- Title: DreamLedger: Where to Refuse World-Model Imagination Using Execution-Settled Credit
- Title(参考訳): DreamLedger: 実行段階のクレジットを使って世界モデルイマジネーションを再利用する方法
- Authors: Xianyao Li, Ruitong Tian, Rui Min, Fang Xu, Jing Du,
- Abstract要約: DreamLedgerは、信頼性を永続的なデプロイオブジェクトとして扱う。
各消費予測はクレームとして登録され、到着した現実に対して解決される。
すべての依存イベントは、依存チケットと再生可能なログを通じて監査可能である。
- 参考スコア(独自算出の注目度): 4.081610854729677
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Robots are beginning to act on world-model predictions, yet reliability is still expressed through instantaneous, model-internal signals that say whether a prediction looks trustworthy now, not where comparable imagination has already failed. DreamLedger instead treats reliability as a persistent deployment object: an execution-settled credit file recording how often consumed predictions are borne out, indexed by operating condition, region, and prediction horizon, and consulted before each use. Each consumed prediction is registered as a claim and settled against arriving reality without manual labels; the resulting credit gates consumption (low credit shortens reliance or triggers observation), and every reliance event remains auditable via dependency tickets and replayable logs. Persistent credit changes where the gate refuses rather than what the model gets wrong: 69% of denials land on cells that have already failed, episode-local reset triples off-target denials in healthy conditions, and under a localized recurrent degradation persistent credit halves burned imagination, at a cost in task completion. Across three simulated domains, unmodified DreamerV3, TD-MPC2, and V-JEPA 2-AC mounts, and a real Franka, paired quadrotor evaluation shows credit-gated planning reduces burned imagination by 62% (95% CI 43-81%) versus blind consumption; settlement-grounded calibration yields moderate, seed-consistent operating points where raw instantaneous gates collapse to extremes, while persistent books trade verification for reliance (manipulation probes 1.00 to 0.36/episode at success 0.98 vs. 0.94). The trust layer spans decoder-, latent-, and token-space interfaces. On hardware, settlement runs under real sensing and contact noise, both models are priced creditworthy at the frozen 9-cm tolerance, a failure loop is re-priced online, and all 1,062 registered spends replay from the audit logs.
- Abstract(参考訳): ロボットは、世界モデル予測に取り組み始めているが、信頼性は、予測がすでに失敗している場所ではなく、現在信頼できるように見えるかどうかを示す、瞬間的なモデル内部信号によって表現されている。
DreamLedgerは、信頼性を永続的なデプロイオブジェクトとして扱う。実行設定のクレジットファイルは、どれだけ頻繁に消費される予測が出力され、操作条件、リージョン、予測水平線によってインデックス付けされ、使用前に相談されるかを記録する。
それぞれの消費予測はクレームとして登録され、手動のラベルなしで到着した現実に対して解決される; 結果として生じるクレジットゲートの消費(クレジットの短縮または観察のトリガー)は、依存チケットと再生可能なログを介して、すべての依存イベントを監査することができる。
69%のデニアルが既に失敗した細胞に着陸し、エピソードローカルリセットは正常な状態でターゲット外デニアルを3倍に減らし、局所的に繰り返し劣化する永続的なクレジットハーフは、タスク完了のコストで想像力を燃やした。
未修正DreamerV3、TD-MPC2、V-JEPA2-ACマウントの3つの模擬ドメインと、実際のFrankaのペアによる2-ACマウントは、クレジット付き計画により、燃えた想像力を62%(95% CI 43-81%)減らす。
信頼層はデコーダ、ラテント、トークン空間のインターフェイスにまたがる。
ハードウェアでは、実際のセンサーとコンタクトノイズ下での決済が実行され、どちらのモデルも凍った9cmの耐久に耐えられる価格で販売され、障害ループはオンラインで再販売され、登録された1,062人は監査ログからリプレイに費やされる。
関連論文リスト
- TRACE: A Multi-Layer Benchmark for Human AI Controller Coordination Under Drift and Failure [0.0]
現代のサイバー物理およびAI支援システムは、人間のオペレータ、AI決定モジュール、自動コントローラを単一の制御ループで結合する。
標準ベンチマークでは、これらのレイヤ間でドリフトと障害がどのように伝播するかの、タイムアラインなマルチレイヤトレースをキャプチャしない。
ALFREDは,日常的な家庭内課題に対する基礎的な指導基準である。
論文 参考訳(メタデータ) (2026-08-07T00:03:55Z) - Zero2Skill: Bootstrapping Robot Skills through Autonomous Data Collection, Training, and Deployment [53.913175773174636]
既存のパイプラインは、自己リセット、VLM検証、言語指導による修正を通じて、人間の労力を減らす。
Zero2Skill(ゼロ2スキル)は、人ロボットの共生型エージェントシステムで、丸ごとの修正を維持・再利用する。
論文 参考訳(メタデータ) (2026-07-15T17:16:24Z) - Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents [58.49879338545782]
ロングホライゾンタスクは、現実のロボット展開では一般的なものだが、そのようなタスクの障害検出は未調査のままである。
動作条件付き世界モデルからの潜在表現を用いて、操作軌跡をモニタする故障検出フレームワークであるForesightを提案する。
この結果から, 動作条件付きワールドモデル埋め込みは, 長距離操作における信頼性のある故障監視のためのスケーラブルな表現を提供する可能性が示唆された。
論文 参考訳(メタデータ) (2026-06-22T09:32:28Z) - Online Reward-Punishment Learning from Fixed-Channel Perceptual Event Streams without Environment Rewards [3.1050763275397806]
OHIRLは4つの役割を分けている: M_psiは次のパケット予測、D_omegaモデルは残留力学、C_etaは固定された内部遷移後軌道である。
C_etaは、回復陽性、永続性/成長陰性残留制御方位を使用する。
条件誤差分解は、B_xiエビデンス推定誤差と残留ポリシー最適化誤差とを分離する。
論文 参考訳(メタデータ) (2026-06-17T11:43:10Z) - Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents [66.97968363332465]
エージェントベンチマークの3つのギャップに対処するエンドツーエンド評価スイートであるClaw-Evalを紹介した。
Claw-Evalは3つのグループにまたがる9つのカテゴリにまたがる300の人間検証タスクで構成されている。
すべてのエージェントアクションは、3つの独立したエビデンスチャネルを通じて記録される。
論文 参考訳(メタデータ) (2026-04-07T17:43:18Z) - Neuro-Symbolic Financial Reasoning via Deterministic Fact Ledgers and Adversarial Low-Latency Hallucination Detector [2.950245545999729]
検証可能な数値推論エージェント(VeNRA)について紹介する。
VeNRAは、RAGパラダイムを確率的テキストの検索から厳密な型付きUniversal Fact Ledger (UFL)による決定論的変数の検索へとシフトさせる
著者らは3ビリオンのSLMを訓練し、単一の推論予算を用いて予測候補を法医学的に監査する。
論文 参考訳(メタデータ) (2026-03-04T22:55:16Z) - Bench2Drive-R: Turning Real World Data into Reactive Closed-Loop Autonomous Driving Benchmark by Generative Model [63.336123527432136]
我々は,リアクティブ閉ループ評価を可能にする生成フレームワークであるBench2Drive-Rを紹介する。
既存の自動運転用ビデオ生成モデルとは異なり、提案された設計はインタラクティブなシミュレーションに適したものである。
我々は、Bench2Drive-Rの生成品質を既存の生成モデルと比較し、最先端の性能を達成する。
論文 参考訳(メタデータ) (2024-12-11T06:35:18Z) - Lazy Layers to Make Fine-Tuned Diffusion Models More Traceable [70.77600345240867]
新たな任意の任意配置(AIAO)戦略は、微調整による除去に耐性を持たせる。
拡散モデルの入力/出力空間のバックドアを設計する既存の手法とは異なり,本手法では,サンプルサブパスの特徴空間にバックドアを埋め込む方法を提案する。
MS-COCO,AFHQ,LSUN,CUB-200,DreamBoothの各データセットに関する実証研究により,AIAOの堅牢性が確認された。
論文 参考訳(メタデータ) (2024-05-01T12:03:39Z) - TRUST-LAPSE: An Explainable and Actionable Mistrust Scoring Framework
for Model Monitoring [4.262769931159288]
連続モデル監視のための"ミストラスト"スコアリングフレームワークであるTRUST-LAPSEを提案する。
我々は,各入力サンプルのモデル予測の信頼性を,潜時空間埋め込みのシーケンスを用いて評価する。
AUROCs 84.1 (vision), 73.9 (audio), 77.1 (clinical EEGs)
論文 参考訳(メタデータ) (2022-07-22T18:32:38Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。