論文の概要: PentestChain: A Cost-Aware, MCP-Orchestrated Framework for Automated Penetration Testing with Free-Tier LLMs
- arxiv url: http://arxiv.org/abs/2609.18120v2
- Date: Fri, 18 Sep 2026 00:47:30 GMT
- ステータス: 翻訳完了
- システム内更新日: 2026-09-21 18:40:16.698709
- Title: PentestChain: A Cost-Aware, MCP-Orchestrated Framework for Automated Penetration Testing with Free-Tier LLMs
- Title(参考訳): PentestChain: フリーティアLCMによる自動貫入テストのための費用対効果を考慮したMPPオーケストレーションフレームワーク
- Abstract要約: PentestChainは10フェーズの自動浸透テストフレームワークである。
コストを意識したAIカスケードと、キュレートされた決定論的エクスプロイトマップを結合する。
現在測定されているレガシーターゲットについて、このフレームワークは26のサービスを検出し、34のCVEを濃縮した。
- 参考スコア(独自算出の注目度): 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: AI-driven penetration testing has been demonstrated with premium frontier models such as GPT-4, but the per-engagement token cost makes continuous, automated testing unaffordable for the smaller organisations that need it most. This paper presents PentestChain, a ten-phase automated penetration testing framework that couples a curated, deterministic exploit map with a cost-aware AI cascade-a local Ollama model (qwen2.5-7b) first, then free-tier OpenRouter and Cerebras, with a rule-based fallback that always produces output-and exposes the full pipeline through a Model Context Protocol (MCP) server with eleven tools. We make three contributions. First, we treat US-dollar cost per engagement as a measured, first-class evaluation metric and show that a 7B-parameter local model, kept off the critical path by a deterministic backbone, sustains end-to-end operation at zero measured paid-API cost. Second, we analyse the attack surface that an MCP-exposed offensive engine introduces, grounding a four-position threat model in the 2025 MCP incident record (the CVE-2025-6514 remote-code-execution flaw in mcp-remote, the postmark-mcp supply-chain backdoor, and the tool-poisoning-rug-pull-line-jumping class), and contribute four mitigations. Third, we specify a reproducible, containerised evalua-tion protocol aligned with the standardised testbeds now expected at top-tier venues-AutoPenBench, a Cybench subset, and the PentestGPT 182-sub-task benchmark-with multi-trial statistics (more than 10 trials per configuration, pass-at-k, non-parametric significance tests and effect sizes) and direct, same testbed reproduction of the PentestGPT and PentestAgent baselines rather than citation of their published numbers. On the legacy targets measured to date, the framework detected 26 services, enriched 34 CVEs, produced
- Abstract(参考訳): AI駆動の浸透テストは、GPT-4のようなプレミアムフロンティアモデルで実証されているが、エンゲージメント単位のトークンコストにより、最も必要な小さな組織では、継続的かつ自動化されたテストが実行不可能になる。
本稿では,PentestChainを提案する。PentestChainは10フェーズの自動貫入テストフレームワークで,コストを意識したAIカスケードとローカルオラマモデル(qwen2.5-7b)を組み合わせ,次にフリーティアのOpenRouterとCerebrasに,出力を常に生成するルールベースのフォールバックと,11つのツールを備えたモデルコンテキストプロトコル(MCP)サーバを通じてパイプライン全体を公開する。
私たちは3つの貢献をします。
まず,7Bパラメータの局所モデルが決定論的バックボーンによって臨界経路を遮断し,測定済みのAPIコストゼロでエンドツーエンドの操作を継続することを示す。
第2に,MPPが排出する攻撃エンジンが導入する攻撃面を分析し,2025年のMPPインシデント記録(mcp-remoteのCVE-2025-6514リモートコード実行欠陥,ポストマーク-mcpサプライチェーンバックドア,ツール-ポゾン-ルーグ-プルラインジャンピングクラス)に4つの脅威モデルを構築し,4つの緩和に寄与する。
第3に,CybenchサブセットであるAutoPenBenchとPentestGPT 182-sub-taskベンチマーク(構成毎の10トライアル,パスアット-k,非パラメトリックな意味テストと効果サイズ)と,PentestGPTおよびPentestAgentベースラインの直接的かつ同一のテストベッド再現を,トップ層で現在期待されているテストベッドと一致させた再現可能なコンテナ化評価プロトコルを規定する。
現在測定されているレガシーターゲットについて、このフレームワークは、34個のCVEを濃縮した26個のサービスを検出した。
関連論文リスト
- Evaluating Frontier AI Agents as Autonomous Clinical Security Auditors [0.0]
本稿では、フロンティアAIエージェントが、構造化された臨床AIセキュリティ監査を自律的に実施できるかどうかをテストする評価タスクを提案する。
事前訓練された臨床予測モデル、患者データセット、手書きの指示が与えられた場合、各エージェントは擬似コードから4つの攻撃を実行する必要がある。
私たちは3つのフロンティアモデルに対して54回の評価を行い、1変種ごとに3回の評価を行った。
論文 参考訳(メタデータ) (2026-07-15T03:21:55Z) - InvestPhilBench: A Multi-Layer Dynamic Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy [0.9924185669370708]
InvestPhilBenchは8つの認知層にまたがる動的ベンチマークである。
v0.6リリースには118のプライマリソース認証投資原則カードが含まれている。
大規模な再現可能なスコア付けには、Benchmark Automated Scoring Pipelineを紹介します。
論文 参考訳(メタデータ) (2026-06-24T15:53:20Z) - CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs [100.38986535324284]
我々は、フロンティアモデル全体でのtextbfcontrol textbfintervention (CI) の認識を測定するベンチマークである textbfCIAware-Bench を紹介する。
CIAware-Benchは、モデルが自身の軌跡を制御介入によって修正されたものと区別できるかどうかをテストする。
論文 参考訳(メタデータ) (2026-06-09T16:24:16Z) - ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox [61.862814740220806]
$textbfComplexMCP$は厳格な条件下でエージェントを評価するために設計されたベンチマークである。
Model Context Protocol (MCP)上に構築された$textbfComplexMCP$は300以上の精巧にテストされたツールを提供する。
論文 参考訳(メタデータ) (2026-05-11T16:20:51Z) - MCP Pitfall Lab: Exposing Developer Pitfalls in MCP Tool Server Security under Multi-Vector Attacks [0.7305019142196584]
MCP Pitfall Labは,開発者の落とし穴を再現可能なシナリオとして運用するプロトコル対応のセキュリティテストフレームワークである。
Pitfall Labは,現実的なマルチベクタ条件下でのMPPツールサーバの実用的,エンドツーエンド評価と強化を可能にする。
論文 参考訳(メタデータ) (2026-04-23T09:39:15Z) - Automated Self-Testing as a Quality Gate: Evidence-Driven Release Management for LLM Applications [51.56484100374058]
我々は,エビデンスに基づくリリース決定を伴う品質ゲートを導入する自動自己テストフレームワークを提案する。
内部展開型多エージェント対話型AIシステムの縦型ケーススタディにより,本フレームワークの評価を行った。
論文 参考訳(メタデータ) (2026-03-13T20:44:15Z) - AgentAssay: Token-Efficient Regression Testing for Non-Deterministic AI Agent Workflows [0.0]
AgentAssayは、非決定論的AIエージェントを回帰テストするための最初のトークン効率のよいフレームワークである。
厳密な統計保証を維持しながら78-100%のコスト削減を実現している。
論文 参考訳(メタデータ) (2026-03-03T04:59:25Z) - Synthesizing File-Level Data for Unit Test Generation with Chain-of-Thoughts via Self-Debugging [40.29934051200609]
本稿では,高品質なUTトレーニングを実現するための新しいデータ蒸留手法を提案する。
このパイプラインをオープンソースプロジェクトの大規模なコーパスに適用します。
実験により, 微調整モデルにより, UT生成効率が高いことを示す。
論文 参考訳(メタデータ) (2026-02-03T06:52:54Z) - AutoPT: How Far Are We from the End2End Automated Web Penetration Testing? [54.65079443902714]
LLMによって駆動されるPSMの原理に基づく自動浸透試験エージェントであるAutoPTを紹介する。
以上の結果から, AutoPT は GPT-4o ミニモデル上でのベースラインフレームワーク ReAct よりも優れていた。
論文 参考訳(メタデータ) (2024-11-02T13:24:30Z) - AI ATAC 1: An Evaluation of Prominent Commercial Malware Detectors [3.0909095595694724]
本研究は,6つの有名な商用エンドポイントマルウェア検出装置,ネットワークマルウェア検出装置,およびサイバー技術ベンダーによるファイル検証アルゴリズムの評価を行う。
この評価は、アメリカ海軍の資金提供を受けたり、完成したりして、AI ATAC(Artificial Intelligence Applications to Autonomous Cybersecurity)賞の1つとして管理された。
論文 参考訳(メタデータ) (2023-08-28T18:46:12Z)
関連論文リストは本サイト内にある論文のタイトル・アブストラクトから自動的に作成しています。
指定された論文の情報です。
本サイトの運営者は本サイト(すべての情報・翻訳含む)の品質を保証せず、本サイト(すべての情報・翻訳含む)を使用して発生したあらゆる結果について一切の責任を負いません。