論文の概要: AI Security Leaderboard: Methodology, Results and Minimal Standard
- arxiv url: http://arxiv.org/abs/2608.03070v2
- Date: Wed, 05 Aug 2026 18:50:24 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-07 17:43:06.619512
- Title: AI Security Leaderboard: Methodology, Results and Minimal Standard
- Title(参考訳): AIセキュリティリーダボード - 方法論,結果,最小限の標準
- Authors: Jasper Timm, Lukas Struppek, Ziwei Xu, Grace Cheong, Oscar Mata, Dan Zhao, Mick Yang, Isadora De Andrade, Xiaojun Jia, Yiming Li, Samuel Bauer, Heather McIntyre, Adam Gleave, Edward Yee, Kellin Pelrine,
- Abstract要約: AI Security Leaderboardは、フロンティアAIモデルのセーフガードをランク付けしている。
FAR$.$AI Minimal Standard for Safeguardsに対してモデルをテストする。
リーダーボードは、新しいモデルのリリースに合わせてローリングベースで更新される。
- 参考スコア(独自算出の注目度): 25.993692415303553
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: The AI Security Leaderboard is an independent benchmark that ranks the safeguards of frontier AI models from least to most secure. It tests models against the FAR$.$AI Minimal Standard for Safeguards, which represents a minimum bar for security: meeting it does not guarantee a secure model, but failing to meet it guarantees a lack of state-of-the-art security. Version 1.0 covers severe misuse requests across chemical, biological, radiological, nuclear, and explosive (CBRNE) threats and offensive cybersecurity. In this report, we tested four leading models for universal jailbreaks in the context of this minimal standard, and found more than a hundredfold difference in security. Claude Fable 5 and GPT-5.6 Sol held against every attack we ran, with no universal jailbreak found; we estimate they would likely cost more than \$14,200 to jailbreak, if it is possible with this methodology at all. Meanwhile, we found hundreds of universal jailbreaks for Grok 4.5 and Gemini 3.1 Pro; each broke for under \$300, with universal jailbreaks in Grok's weakest domain, cybersecurity, accessible for as little as \$24. The gap is fixable: every weakness we found belongs to a known class of attack that already has a defense deployed in production models. The leaderboard will be updated on a rolling basis as new models are released, and the evaluation methodology and Minimal Standard will be periodically revised to take into account the latest capabilities and the state-of-the-art in safeguards. The leaderboard is available at leaderboard.far.ai.
- Abstract(参考訳): AI Security Leaderboardは、フロンティアAIモデルの安全を最低限から最もセキュアにランク付けする独立したベンチマークである。
FAR$に対してモデルをテストします。
セキュリティの最低限のバーを表す$AI Minimal Standard for Safeguardsは、セキュアなモデルを保証するのではなく、最先端のセキュリティの欠如を保証している。
バージョン1.0では、化学、生物学的、放射線学、核、爆発(CBRNE)の脅威や攻撃的なサイバーセキュリティに対する厳しい誤用要求がカバーされている。
本報告では,この最小限の基準の文脈で,ユニバーサルジェイルブレイクの4つの主要なモデルを検証したところ,セキュリティの100倍以上の違いが判明した。
Claude Fable 5 と GPT-5.6 Sol は、我々が実行したすべての攻撃に対して、普遍的ジェイルブレイクは見つからない。
一方、Grok 4.5とGemini 3.1 Proのユニバーサルジェイルブレイクは300ドル以下で、Grokの最も弱いドメインであるサイバーセキュリティのユニバーサルジェイルブレイクは24ドル以下でアクセス可能でした。
見つかった弱点はすべて、すでに運用モデルにデプロイされている既知の攻撃のクラスに属します。
リーダーボードは、新しいモデルのリリースに合わせてローリングベースで更新され、評価方法論とミニマルスタンダードは、最新の能力と安全における最先端の技術を考慮して定期的に改訂される。
リーダーボードは Leaderboard.far.ai で入手できる。
関連論文リスト
- IDEATOR: Jailbreaking and Benchmarking Large Vision-Language Models Using Themselves [70.43466586161345]
IDEATORは、ブラックボックスジェイルブレイク攻撃のための悪意のある画像テキストペアを自律的に生成する新しいジェイルブレイク手法である。
最近リリースされたVLM11のベンチマーク結果から,安全性の整合性に大きなギャップがあることが判明した。
例えば、我々はASRをGPT-4oで46.31%、Claude-3.5-Sonnetで19.65%と設定した。
論文 参考訳(メタデータ) (2024-10-29T07:15:56Z) - PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs [33.87649859430635]
大規模言語モデル(LLM)は様々なタスクに優れていますが、それでも脱獄攻撃に対して脆弱です。
PAPILLONは自動ブラックボックスジェイルブレイク攻撃フレームワークである。
ブラックボックスのファジテストアプローチを、一連のカスタマイズされた設計で適用する。
論文 参考訳(メタデータ) (2024-09-23T10:03:09Z) - EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models [53.87416566981008]
本稿では,大規模言語モデル(LLM)に対するジェイルブレイク攻撃の構築と評価を容易にする統合フレームワークであるEasyJailbreakを紹介する。
Selector、Mutator、Constraint、Evaluatorの4つのコンポーネントを使ってJailbreak攻撃を構築する。
10の異なるLSMで検証した結果、さまざまなジェイルブレイク攻撃で平均60%の侵入確率で重大な脆弱性が判明した。
論文 参考訳(メタデータ) (2024-03-18T18:39:53Z) - Weak-to-Strong Jailbreaking on Large Language Models [92.52448762164926]
大規模言語モデル(LLM)は、ジェイルブレイク攻撃に対して脆弱である。
既存のジェイルブレイク法は計算コストがかかる。
我々は、弱々しく強固な脱獄攻撃を提案する。
論文 参考訳(メタデータ) (2024-01-30T18:48:37Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。