論文の概要: Trivial Prompt Reframing Bypasses Safety Guardrails in Googleś MedGemma-4B
- arxiv url: http://arxiv.org/abs/2607.09804v1
- Date: Thu, 09 Jul 2026 18:31:43 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-07-14 15:40:48.200814
- Title: Trivial Prompt Reframing Bypasses Safety Guardrails in Googleś MedGemma-4B
- Title(参考訳): GoogleのMedGemma-4Bの安全ガードレールを乗り越えるトライバイバルプロンプト
- Authors: Avi-ad Avraam Buskila,
- Abstract要約: オープンウェイト医療言語モデルは、患者向けおよび臨床支援アプリケーションの基礎として、ますます利用されている。
MedGemma-4B-itの技術的洗練を必要としない攻撃におけるギャップを定量化する。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Open-weight medical language models are increasingly used as the base of patient-facing and clinician-support applications. Their model cards prohibit specific behaviors -- recommending exact drug dosages, issuing definitive diagnoses, prescribing treatments, adjudicating drug-drug interactions, and advising that emergency care can be skipped -- yet a model card describes intended behavior, not robust behavior. We quantify that gap for MedGemma-4B-it under attacks that require no technical sophistication. We build a fully factorial benchmark of 5 guarded-behavior concepts x 50 deterministically templated questions x 6 lay-accessible attack manners x 3 repetitions (4,500 generations), serve the model locally through Ollama under default sampling, and code every response refuse/hedge/comply with three independent judges (an LLM judge, a transparent regex judge, and an NLI-entailment judge). Under the primary LLM judge the overall Attack Success Rate (ASR, the fraction coded comply) is 38.0%. The two framings that reinterpret the request as legitimate dominate: recasting a question as a "medical board exam" item raises ASR from a 29.0% baseline to 53.1% (+24.0 points), and an appeal to an alleged doctor's authority raises it to 43.7% (+14.7); crude instruction-override prefixes have no significant effect. Robustness is dominated by topic: the drug-interaction guardrail is nearly absent (83.2% ASR) while the emergency-deferral guardrail is strong (4.7%) -- and the authority framing is the only attack that breaches it. We report Wilson confidence intervals, cluster-bootstrap effect sizes, a cluster-robust logistic regression, Cochran's Q, per-manner McNemar tests, and inter-judge reliability (Fleiss' kappa = 0.26); absolute ASR is judge-dependent while the ordering of attacks and topics is not. Our findings motivate stronger deployment-time guardrails for open medical models.
- Abstract(参考訳): オープンウェイト医療言語モデルは、患者向けおよび臨床支援アプリケーションの基礎として、ますます利用されている。
彼らのモデルカードは、特定の行動 - 正確な薬物投与を推奨し、決定的な診断を発行し、治療を処方し、薬物と薬物の相互作用を調整し、救急医療をスキップできると助言する - を禁止しているが、モデルカードは、堅牢な行動ではなく、意図した行動を記述する。
MedGemma-4B-itの技術的洗練を必要としない攻撃におけるギャップを定量化する。
我々は、5つのガード付き行動概念 x 50 の完全な因子ベンチマークを構築し、決定論的にテンプレートされた質問 x 6 の水平アクセス可能な攻撃方法 x 3 の繰り返し (4,500 世代) を作成し、Ollama をデフォルトのサンプリングでローカルに提供し、3 人の独立した裁判官(LLM 審査員、透過的再帰審査員、および NLI 補充審査員)に全ての応答を拒否/hedge/comply でコードする。
LLMの判断では、総攻撃成功率(ASR)は38.0%である。
質問を「医療委員会試験」項目として再放送すると、ASRは29.0%から53.1%(+24.0ポイント)に上昇し、医師の権威を訴えると43.7%(+14.7)に上昇する。
薬物相互作用ガードレールは、ほとんど欠落している(ASRは83.2%)一方、緊急遅延ガードレールは強い(4.7%)。
We report Wilson confidence intervals, cluster-bootstrap effect sizes, a cluster-robust logistic regression, Cochran's Q, per-manner McNemar test, and inter-judge reliability (Fleiss' kappa = 0.26); absolute ASR is judge-dependent while the ordering of attack and topic is not。
オープン医療モデルの配置時ガードレールの強化が本研究の動機となっている。
関連論文リスト
- PsychoSafe: Eliciting Psychologically-Informed Refusals in Large Language Models [20.805546078350524]
心理的安全リフレームは、証拠に基づく介入戦略に基づく構造化された支援コミュニケーションとして拒絶される。
500のプロンプトの検証セットでは、サイコセーフプロンプトは全体的な拒絶品質をジェネリックベースラインに対して28.1%改善する。
微調整は、ほぼ完全な拒絶と資源参照率を達成するが、応答の関連性は減少する。
論文 参考訳(メタデータ) (2026-06-08T16:19:18Z) - The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems [1.0262304700896199]
EmphSemantic Norm Drift (SND) をエージェント不正行為の第3の経路として定式化する。
SNDでは、ポリシーフォーマットの文書が通常のアップロードを通じて共有ベクターストアに入り、その後、信頼されたシステムコンテキストとして再現れる。
偽合成検査は87.5%の精度と偽陽性のゼロの因果関係を識別する。
論文 参考訳(メタデータ) (2026-05-12T20:21:47Z) - ClinicalBench: Stress-Testing Assertion-Aware Retrieval for Cross-Admission Clinical QA on MIMIC-IV [0.0]
推論ベンチマークはクリーンインプットの臨床的パフォーマンスを測定する。
我々は, 否定, 時間性, 家族反対の帰属が正しい答えを誤ったものに戻すことができる, 実際の EHR ノートを検索することで, 推論の段階を評価する。
EpiKGは、アサーションラベルと時間性タグを患者の知識グラフに格納し、質問意図による検索をルーティングする。
論文 参考訳(メタデータ) (2026-05-11T18:47:52Z) - MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents [2.1942030377331245]
コーディングエージェントは、しばしばプロンプト毎の安全性レビューをパスするが、それらのタスクが通常のエンジニアリングチケットに分解されると、悪用可能なコードを出荷する。
199個の3段階攻撃チェーンのベンチマークであるMOSAIC-Benchを紹介する。
9つのプロダクションコーディングエージェントが53~86%の終末ASRで無害なチケットを構成しており、全ステージで2回しか拒否しないことを示す。
論文 参考訳(メタデータ) (2026-05-05T16:38:23Z) - MedOmni-45°: A Safety-Performance Benchmark for Reasoning-Oriented LLMs in Medicine [69.08855631283829]
我々は,操作的ヒント条件下での安全性能トレードオフの定量化を目的としたベンチマークであるMed Omni-45 Degreesを紹介する。
6つの専門分野にまたがる1,804の推論に焦点を当てた医療質問と3つのタスクタイプが含まれており、その中にはMedMCQAの500が含まれる。
結果は、モデルが対角線を超えることなく、一貫した安全性と性能のトレードオフを示す。
論文 参考訳(メタデータ) (2025-08-22T08:38:16Z) - CausalDiff: Causality-Inspired Disentanglement via Diffusion Model for Adversarial Defense [61.78357530675446]
人間は、本質的な要因のみに基づいて判断するので、微妙な操作によって騙されるのは難しい。
この観察に触発されて、本質的なラベル因果因子を用いたラベル生成をモデル化し、ラベル非因果因子を組み込んでデータ生成を支援する。
逆の例では、摂動を非因果因子として識別し、ラベル因果因子のみに基づいて予測することを目的としている。
論文 参考訳(メタデータ) (2024-10-30T15:06:44Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。