論文の概要: Structured Visual Narratives Undermine Safety Alignment in Multimodal Large Language Models
- arxiv url: http://arxiv.org/abs/2603.21697v1
- Date: Mon, 23 Mar 2026 08:32:09 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-03-24 19:11:39.569257
- Title: Structured Visual Narratives Undermine Safety Alignment in Multimodal Large Language Models
- Title(参考訳): マルチモーダル大言語モデルにおける安全アライメントを損なう構造的ビジュアルナラティブ
- Authors: Rui Yang Tan, Yujia Hu, Roy Ka-Wei Lee,
- Abstract要約: 単純な3パネルの視覚的物語の中に有害な目標を埋め込むコミック・テンポレート・ジェイルブレイクについて研究する。
ComicJailbreakは、コミックベースのジェイルブレイクベンチマークであり、10の有害カテゴリと5つのタスク設定にまたがる1,167の攻撃インスタンスがある。
- 参考スコア(独自算出の注目度): 12.740730240000468
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Multimodal Large Language Models (MLLMs) extend text-only LLMs with visual reasoning, but also introduce new safety failure modes under visually grounded instructions. We study comic-template jailbreaks that embed harmful goals inside simple three-panel visual narratives and prompt the model to role-play and "complete the comic." Building on JailbreakBench and JailbreakV, we introduce ComicJailbreak, a comic-based jailbreak benchmark with 1,167 attack instances spanning 10 harm categories and 5 task setups. Across 15 state-of-the-art MLLMs (six commercial and nine open-source), comic-based attacks achieve success rates comparable to strong rule-based jailbreaks and substantially outperform plain-text and random-image baselines, with ensemble success rates exceeding 90% on several commercial models. Then, with the existing defense methodologies, we show that these methods are effective against the harmful comics, they will induce a high refusal rate when prompted with benign prompts. Finally, using automatic judging and targeted human evaluation, we show that current safety evaluators can be unreliable on sensitive but non-harmful content. Our findings highlight the need for safety alignment robust to narrative-driven multimodal jailbreaks.
- Abstract(参考訳): MLLM(Multimodal Large Language Models)は、テキストのみのLLMを視覚的推論で拡張すると同時に、視覚的に接地された命令の下で新しい安全性障害モードを導入する。
本研究では,単純な3パネルの視覚的物語の中に有害な目標を埋め込んだ喜劇的ジェイルブレイクについて検討し,そのモデルにロールプレイと「コミックを完成させる」よう促す。
JailbreakBenchとJailbreakVをベースに構築されたComicJailbreakは、コミックベースのジェイルブレイクベンチマークであり、10の有害カテゴリと5つのタスク設定にまたがる1,167の攻撃インスタンスがある。
15の最先端のMLLM(6つの商用および9つのオープンソース)で、コミックベースの攻撃は強力なルールベースのジェイルブレイクに匹敵する成功率を獲得し、いくつかの商業モデルにおいてアンサンブルの成功率が90%を超えるようなプレーンテキストとランダムイメージのベースラインを大幅に上回っている。
そして,既存の防衛手法を用いて,これらの手法が有害漫画に対して有効であることを示す。
最後に, 自動判定と人体評価を用いて, 現行の安全性評価装置は, センシティブな内容と無害な内容で信頼性が低いことを示す。
本研究は,物語駆動型マルチモーダルジェイルブレイクに対して,安全アライメントの必要性を強調した。
関連論文リスト
- Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling [11.939828002077482]
MLLM(Multimodal large language model)は、優れた能力を示すが、ジェイルブレイク攻撃の影響を受けない。
本研究では,最新のMLLMにおける安全アライメントを回避するために,連続的な漫画スタイルの視覚的物語を活用する新しい手法を提案する。
攻撃成功率は平均83.5%であり, 先行技術の46%を突破した。
論文 参考訳(メタデータ) (2025-10-16T18:30:26Z) - IDEATOR: Jailbreaking and Benchmarking Large Vision-Language Models Using Themselves [70.43466586161345]
IDEATORは、ブラックボックスジェイルブレイク攻撃のための悪意のある画像テキストペアを自律的に生成する新しいジェイルブレイク手法である。
最近リリースされたVLM11のベンチマーク結果から,安全性の整合性に大きなギャップがあることが判明した。
例えば、我々はASRをGPT-4oで46.31%、Claude-3.5-Sonnetで19.65%と設定した。
論文 参考訳(メタデータ) (2024-10-29T07:15:56Z) - Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation [71.92055093709924]
そこで本稿では, ガーブレッドの逆数プロンプトを, 一貫性のある, 可読性のある自然言語の逆数プロンプトに"翻訳"する手法を提案する。
また、jailbreakプロンプトの効果的な設計を発見し、jailbreak攻撃の理解を深めるための新しいアプローチも提供する。
本稿では,AdvBench上でのLlama-2-Chatモデルに対する攻撃成功率は90%以上である。
論文 参考訳(メタデータ) (2024-10-15T06:31:04Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。