論文の概要: Understanding as an Explicit and Assessable Component of Frontier AI Safety Decisions
- arxiv url: http://arxiv.org/abs/2608.19816v2
- Date: Fri, 21 Aug 2026 10:28:50 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-24 14:49:31.95054
- Title: Understanding as an Explicit and Assessable Component of Frontier AI Safety Decisions
- Title(参考訳): フロンティアAIの安全性決定の明示的で評価可能なコンポーネントとしての理解
- Abstract要約: 意思決定者は、トレーニングやフロンティアAIシステムのデプロイについて、十分な理解を必要とする。
本稿では,Elgin と Arendt からの理解の哲学的基盤を運用するための Assurance 2.0 フレームワークを用いた安全事例の最近の展開について述べる。
- 参考スコア(独自算出の注目度): 35.25345433558124
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Decision makers need sufficient understanding to make good decisions about training or deploying frontier AI systems. However, such decisions are increasingly made under time-pressure, and this combined with the use of AI generated artefact creation, can mean that the existence of safety cases and system cards may no longer demonstrate that sufficient understanding exists. Our provisional methodology for making understanding explicit and assessable requires the production of an explicit description of 4 objects of understanding (decision, decision-frame, safety justification, system-in-context) and a justification for the adequacy of this understanding. In addition, the methodology provides a mechanism for describing and evaluating the adequacy of the decision-maker representation of this understanding. It builds on recent developments in safety cases using the Assurance 2.0 framework to operationalise the philosophical basis of understanding from Elgin and Arendt. To assess the methodology we trialled two different scenarios. One scenario, which we investigated through role-based analysis, concerned the risk of scheming in the deployment of an AI coding agent in a robotics company and the other scenario was for the higher uncertainty, more decision-critical argument of 'If Anyone Builds It, Everyone Dies' (Yudkowsky and Soares). The trial's central finding, for these two scenarios, is that the methodology could be applied and was found to be generative: we found the analyses that justify sufficiency of understanding (internal coherence, tethering, felicitous falsehoods, external coherence) drives the engineering.
- Abstract(参考訳): 意思決定者は、トレーニングやフロンティアAIシステムのデプロイについて、十分な理解を必要とする。
しかし、こうした決定は時間的プレッシャーの下でますます行われており、AIが生成した人工物の生成と組み合わせることで、安全ケースとシステムカードの存在は、もはや十分な理解が存在しないことを証明できない可能性がある。
理解を明確かつ評価可能なものにするには,4つの理解対象(決定,意思決定,安全正当性,システム・イン・コンテクスト)の明示的な記述と,この理解の妥当性の正当化が必要である。
さらに、この方法論は、この理解の意思決定者表現の妥当性を記述し、評価するためのメカニズムを提供する。
Assurance 2.0フレームワークを使用して、ElginとArendtの哲学的理解の基盤を運用している。
方法論を評価するために、我々は2つの異なるシナリオを試した。
ひとつは、ロボット企業におけるAIコーディングエージェントのデプロイにおけるスケジューリングのリスクについて、もうひとつは、"If Anyone Builds It, Everyone Dies"(Yudkowsky氏とSoares氏)のより不確実で、より決定的に批判的な議論である。
この2つのシナリオにおいて、トライアルの中心的な発見は、この方法論が適用可能であり、生成可能であることが判明したことである。
関連論文リスト
- Understanding: reframing automation and assurance [0.0]
安全と保証のケースは、責任あるエンジニアリングとガバナンスの決定に必要な理解から切り離されるリスクがあります。
理解は明確で、評価可能で、防御可能な意思決定要素になるべきだ、と私たちは主張する。
論文 参考訳(メタデータ) (2026-04-07T10:04:26Z) - A Survey of Reasoning in Autonomous Driving Systems: Open Challenges and Emerging Paradigms [49.66022971508878]
私たちは、推論はモジュラーコンポーネントからシステムの認知コアに高めるべきだと論じています。
応答性推論のトレードオフやソーシャルゲーム推論など,7つの中核的推論課題を導出し,体系化する。
我々は,LLMに基づく推論と,ミリ秒スケールで安全クリティカルな車両制御の要求との間の,高レイテンシ,熟考的特性の根本的かつ未解決な緊張関係を同定する。
論文 参考訳(メタデータ) (2026-03-11T07:40:53Z) - AI Deception: Risks, Dynamics, and Controls [153.71048309527225]
このプロジェクトは、AI偽装分野の包括的で最新の概要を提供する。
我々は、動物の偽装の研究からシグナル伝達理論に基づく、AI偽装の正式な定義を同定する。
我々は,AI偽装研究の展望を,偽装発生と偽装処理の2つの主要な構成要素からなる偽装サイクルとして整理する。
論文 参考訳(メタデータ) (2025-11-27T16:56:04Z) - Too Much to Trust? Measuring the Security and Cognitive Impacts of Explainability in AI-Driven SOCs [0.6990493129893112]
説明可能なAI(XAI)は、AIによる脅威検出の透明性と信頼性を高めるための大きな約束を持っている。
本研究は、セキュリティコンテキストにおける現在の説明手法を再評価し、SOCに適合したロールアウェアでコンテキストに富んだXAI設計が実用性を大幅に向上できることを実証する。
論文 参考訳(メタデータ) (2025-03-03T21:39:15Z) - Combining AI Control Systems and Human Decision Support via Robustness and Criticality [53.10194953873209]
我々は、逆説(AE)の方法論を最先端の強化学習フレームワークに拡張する。
学習したAI制御システムは、敵のタンパリングに対する堅牢性を示す。
トレーニング/学習フレームワークでは、この技術は人間のインタラクションを通じてAIの決定と説明の両方を改善することができる。
論文 参考訳(メタデータ) (2024-07-03T15:38:57Z) - An Objective Metric for Explainable AI: How and Why to Estimate the
Degree of Explainability [3.04585143845864]
本稿では, 客観的手法を用いて, 正しい情報のeX説明可能性の度合いを測定するための, モデルに依存しない新しい指標を提案する。
私たちは、医療とファイナンスのための2つの現実的なAIベースのシステムについて、いくつかの実験とユーザースタディを設計しました。
論文 参考訳(メタデータ) (2021-09-11T17:44:13Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。