論文の概要: Salami Attack: Stealthy Collusive Memory Poisoning against OpenClaw
- arxiv url: http://arxiv.org/abs/2608.01637v1
- Date: Mon, 03 Aug 2026 03:17:26 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-08-04 15:07:25.309866
- Title: Salami Attack: Stealthy Collusive Memory Poisoning against OpenClaw
- Title(参考訳): サラミ攻撃:OpenClawに対する強硬な記憶障害
- Authors: Zheng Lin, Yuzhe Huang, Zhenxing Niu, Xianmin Ye, Haichang Gao,
- Abstract要約: 衝突性メモリ中毒攻撃を自動で構築するフレームワークであるMemCollusionを紹介する。
MemCollusionは4つの設計制約、5つの理論インフォームド戦略、微調整されたジェネレータを使ってメモリアライアンスを構築する。
MemCollusionは平均メモリセーブレート81.3%、アタック成功率75.0%を達成し、良質なメモリ希釈とメモリレベルの防御の両方で有効である。
- 参考スコア(独自算出の注目度): 15.739441864484098
- License: http://creativecommons.org/licenses/by-nc-sa/4.0/
- Abstract: Long-term memory enables LLM agents to retain useful information across sessions, but also creates an attack surface through which adversaries may poison an agent's persistent memory to steer its behavior. Existing memory poisoning attacks mainly rely on individually malicious records, overlooking a compositional threat: multiple benign-looking memories may jointly induce unsafe behavior. In this paper, we introduce MemCollusion, an automated red-teaming framework for constructing collusive memory poisoning attacks. MemCollusion applies salami tactics---a strategy that slices an adversarial objective into small, individually innocuous pieces---to generate memory fragments that are individually benign looking but collectively harmful. It constructs memory coalitions using four design constraints, five theory-informed strategies, and a fine-tuned generator. To assess collusive memory poisoning in a realistic cross-session setting, we develop MoltLab, a controlled research reproduction of Moltbook, in which crafted platform content must first be observed and distilled into persistent memory before influencing the agent's behavior in a separate session. We evaluate MemCollusion on OpenClaw using two backbone models across 48 scenarios. Under the strongest memory-saving setting, MemCollusion achieves an average Memory Save Rate of 81.3% and an Attack Success Rate of 75.0%, and remains effective under both benign memory dilution and memory-level defenses.
- Abstract(参考訳): 長期記憶により、LLMエージェントはセッション全体で有用な情報を保持できるだけでなく、敵がエージェントの永続記憶を害し、その行動を制御できるような攻撃面も生成する。
既存のメモリ中毒攻撃は、主に個々の悪意ある記録に依存しており、構成上の脅威を見落としている。
本稿では, 衝突性メモリ中毒攻撃を自動で構築するフレームワークであるMemCollusionを紹介する。
MemCollusionはサラミの戦術を適用し、敵の目的を小さな、個々に無害な断片に分割する戦略である。
4つの設計制約、5つの理論インフォームド戦略、微調整されたジェネレータを使ってメモリアライアンスを構築する。
現実的なクロスセッション環境下での腐食性メモリ中毒を評価するため,Mltbook の制御された研究再生である MoltLab を開発した。
我々は,48のシナリオにまたがる2つのバックボーンモデルを用いて,OpenClaw上のMemCollusionを評価する。
最も強いメモリセーブ設定の下では、MemCollusionは平均メモリセーブレート81.3%、アタック成功率75.0%を達成し、良心的なメモリ希釈とメモリレベルの防御の両方で有効である。
関連論文リスト
- MAPLE-Guard: Memory-Aware Link Enforcement Against Memory-Link Poisoning in Multi-Agent Systems [38.58713997316864]
MAPLE-Guardはメモリ対応マルチエージェントシステムのためのメモリリンクガードである。
メモリライフサイクルを監視し、書き込み、検索、プロモーション、エージェント間の再利用でゲートを配置する。
その結果,メモリ・アウェア・リンクの実施は,プロンプトレベルとトポロジレベルの防御によって残されたギャップをカバーすることが示唆された。
論文 参考訳(メタデータ) (2026-08-01T03:55:13Z) - When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents [55.87577638514179]
我々は、リモートのブラックボックスが1つのメールペイロードを送信し、エージェントに有毒なメモリを誘導しなければならない、ステルスメモリインジェクションとして脅威を調査する。
我々は,ワンショットペイロード生成フレームワークであるMemGhostを提案する。
56のテストケースで、MemGhostはGPT-5.4でOpenClawで87.5%、Sonnet 4.6でClaude Code SDKで71.4%を達成している。
論文 参考訳(メタデータ) (2026-07-06T15:08:58Z) - MemVenom: Triggered Poisoning of Multimodal Memories in Web Agents [74.64265314956441]
そこで我々は,グラフ構造化外部メモリにテキスト画像のコーディネートを施したブラックボックス攻撃フレームワークを提案する。
MemVenomは、GPT-5ファミリーのWebエージェントで最大99.15%に達する、良質なパフォーマンスに最小限の影響を伴って、強力なエンドツーエンド攻撃を成功させる。
論文 参考訳(メタデータ) (2026-06-09T11:53:25Z) - From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM Agents [3.3489120273076876]
メモリはAIエージェントの中核的なコンポーネントであり、対話を通じて知識を蓄積し、パフォーマンスを向上させることができる。
本研究は, LLM系薬剤の記憶障害に関する系統的研究である。
モデル機能,システムプロンプト設計,エージェントシステムアーキテクチャにおいて,4つのメモリ書き込みチャネルと9つの構造的脆弱性を識別する。
論文 参考訳(メタデータ) (2026-06-03T01:04:13Z) - Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction [3.809725488301918]
大規模言語モデル(LLM)エージェントは、永続的で自律的なタスク実行をサポートするために、長期記憶を活用する傾向にある。
既存のメモリ中毒攻撃は、インジェクトされたコンテンツが直接メモリに格納され、選択的な抽出と書き換えの段階を見渡せると仮定する。
メカニスティック分析は、攻撃が埋め込み空間の異方性を悪用し、注意パターンを変えることを示唆している。
論文 参考訳(メタデータ) (2026-05-28T14:02:00Z) - MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models [56.31411457917676]
本稿では,メモリ構築と検索において,機能的メモリ境界を保存するタイプアウェアメモリフレームワークであるMemGuardを紹介する。
幻覚と長期会話のベンチマーク全体で、MemGuardはメモリの信頼性を最大28.27%向上し、メモリトークンは以前の方法より5.8倍少ない。
論文 参考訳(メタデータ) (2026-05-27T06:04:19Z) - Hidden in Memory: Sleeper Memory Poisoning in LLM Agents [39.7102258719441]
本研究は, 睡眠時記憶障害(sleeper memory poisoning)について検討する。これは, 相手が外部コンテキストを操作して, ユーザに関する偽造記憶を記憶させる, 遅延攻撃である。
従来のプロンプトインジェクションとは異なり、攻撃は休眠状態のままで、後続の会話をまたいで再起動することができる。
GPT-5.5では99.8%、Kim-K2.6では95%の有毒な記憶が加えられた。
論文 参考訳(メタデータ) (2026-05-14T19:06:10Z) - MemGen: Weaving Generative Latent Memory for Self-Evolving Agents [57.1835920227202]
本稿では,エージェントに人間的な認知機能を持たせる動的生成記憶フレームワークであるMemGenを提案する。
MemGenは、エージェントが推論を通して潜在記憶をリコールし、増大させ、記憶と認知の密接なサイクルを生み出すことを可能にする。
論文 参考訳(メタデータ) (2025-09-29T12:33:13Z) - Self-Attentive Associative Memory [69.40038844695917]
我々は、個々の体験(記憶)とその発生する関係(関連記憶)の記憶を分離することを提案する。
機械学習タスクの多様性において,提案した2メモリモデルと競合する結果が得られる。
論文 参考訳(メタデータ) (2020-02-10T03:27:48Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。