Fugu-MT 論文翻訳(概要): From Clouds to Hallucinations: Atmospheric Retrieval Hijacking in Remote Sensing Vision-Language RAG

論文の概要: From Clouds to Hallucinations: Atmospheric Retrieval Hijacking in Remote Sensing Vision-Language RAG

arxiv url: http://arxiv.org/abs/2605.07273v1
Date: Fri, 08 May 2026 05:36:23 GMT
ステータス: 翻訳完了
システム内更新日: 2026-05-11 19:43:38.826126
Title: From Clouds to Hallucinations: Atmospheric Retrieval Hijacking in Remote Sensing Vision-Language RAG
Title（参考訳）: 雲から幻覚へ:リモートセンシングビジョンランゲージRAGにおける大気検索ハイジャック
Authors: Jiaju Han, Chao Li, Chengyin Hu, Qike Zhang, Xuemeng Sun, Xin Wang, Fengyu Zhang, Xiang Chen, Yiwei Wei, Jiahuan Long, Jiujiang Guo,
Abstract要約: CloudWebは、入力イメージのみを修正しつつ、レトリバー、ジェネレータ、知識ベースをデプロイ時に固定したままにしておく、大気検索ハイジャック攻撃である。我々は、GeoRSCLIP、RemoteCLIP、OpenAI CLIP、OpenCLIPを含む5つのCLIPスタイルレトリバーを備えた7データセットリモートセンシングRAGベンチマークでCloudWebを評価した。 CloudWebは、レトリバー全体にわたって、クリーンな検索、手作りの大気ベースライン、ランダムな雲の摂動、そして気象関連の証拠をトップランクの結果に注入する固定された変種を一貫して上回っている。
参考スコア（独自算出の注目度）: 12.942958995976307
License: http://creativecommons.org/licenses/by/4.0/
Abstract: Multimodal RAG systems increasingly rely on vision-language retrievers to ground visual queries in external textual evidence. Existing adversarial studies on RAG mainly manipulate the retrieval corpus or memory, while attacks on vision-language and remote sensing models typically target end-task predictions. Input-space threats to the evidence retrieval stage of remote sensing multimodal RAG remain underexplored. To address this gap, we introduce CloudWeb, an atmospheric retrieval hijacking attack that modifies only the input image while keeping the retriever, generator, and knowledge base fixed at deployment. CloudWeb overlays parameterized cloud- and haze-like patterns on remote sensing images and optimizes them with a retrieval-oriented objective that pulls adversarial image embeddings toward target atmospheric evidence, suppresses source-scene evidence, enforces rank separation, and regularizes naturalness and coverage. To the best of our knowledge, this is the first study of retrieval-stage atmospheric evidence hijacking in remote sensing multimodal RAG. We evaluate CloudWeb on a seven-dataset remote sensing RAG benchmark with five CLIP-style retrievers, including GeoRSCLIP, RemoteCLIP, OpenAI CLIP, and OpenCLIP, together with downstream vision-language generators. Across retrievers, CloudWeb consistently outperforms clean retrieval, handcrafted atmospheric baselines, random cloud perturbations, and fixed variants in injecting weather-related evidence into top-ranked results. On GeoRSCLIP ViT-B/32, Weather@5 increases from 0.71\% to 43.29\%. Downstream generation further shows measurable weather hallucination and semantic shift, indicating that retrieval-stage hijacking can propagate to the final RAG response. These findings reveal a practical failure mode: natural-looking atmospheric changes can compromise evidence retrieval before generation begins.
Abstract（参考訳）: マルチモーダルRAGシステムは、視覚的なクエリを外部のテキストエビデンスでグラウンド化するために、視覚言語レトリバーにますます依存している。 RAGの既存の敵研究は、主に検索コーパスやメモリを操作するが、視覚言語やリモートセンシングモデルに対する攻撃は通常、エンドタスク予測をターゲットとしている。リモートセンシングマルチモーダルRAGのエビデンス検索段階への入力空間の脅威は未解明のままである。このギャップに対処するために、我々はCloudWebを紹介します。これは、入力画像だけを変更できる大気検索ハイジャック攻撃で、レトリバー、ジェネレータ、知識ベースをデプロイ時に固定しつつ、入力画像だけを修正します。 CloudWebは、パラメータ化されたクラウドやヘイズのようなパターンをリモートセンシングイメージ上にオーバーレイし、ターゲットの大気証拠に向けて敵画像の埋め込みを引っ張り、ソースシーンの証拠を抑圧し、ランク分離を強制し、自然さとカバレッジを規則化する、検索指向の目的でそれらを最適化する。我々の知る限りでは、リモートセンシングマルチモーダルRAGにおける検索段階の大気証拠ハイジャックに関する最初の研究である。我々は、GeoRSCLIP、RemoteCLIP、OpenAI CLIP、OpenCLIPを含む5つのCLIPスタイルレトリバーと、下流の視覚言語ジェネレータを備えた7データセットのリモートセンシングRAGベンチマークでCloudWebを評価した。 CloudWebは、レトリバー全体にわたって、クリーンな検索、手作りの大気ベースライン、ランダムな雲の摂動、そして気象関連の証拠をトップランクの結果に注入する固定された変種を一貫して上回っている。 GeoRSCLIP ViT-B/32 では、Weather@5 は 0.71 % から 43.29 % に増加する。下流世代はさらに、観測可能な気象幻覚とセマンティックシフトを示し、検索段階のハイジャックが最終的なRAG応答に伝播することを示した。自然に見える大気の変化は、生成が始まる前に証拠の検索を損なう可能性がある。

論文の概要: From Clouds to Hallucinations: Atmospheric Retrieval Hijacking in Remote Sensing Vision-Language RAG

関連論文リスト