Fugu-MT 論文翻訳(概要): Runtime Execution Traces Guided Automated Program Repair with Multi-Agent Debate

論文の概要: Runtime Execution Traces Guided Automated Program Repair with Multi-Agent Debate

arxiv url: http://arxiv.org/abs/2604.02647v1
Date: Fri, 03 Apr 2026 02:23:25 GMT
ステータス: 翻訳完了
システム内更新日: 2026-04-06 17:20:24.28112
Title: Runtime Execution Traces Guided Automated Program Repair with Multi-Agent Debate
Title（参考訳）: マルチエージェントディベートによる自動プログラム修正をガイドした実行時実行トレース
Authors: Jiaqing Wu, Tong Wu, Manqing Zhang, Yunwei Dong, Bo Shen,
Abstract要約: 自動プログラム修復(APR)は複雑なロジックエラーとサイレント障害に悩まされる。現在のLLMベースのAPRメソッドは主に静的であり、ソースコードと基本的なテスト出力に依存している。我々は、パッチ検証のための共有制約としてランタイム事実を活用するマルチエージェントフレームワークであるTraceRepairを提案する。
参考スコア（独自算出の注目度）: 8.424102114588559
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Abstract: Automated Program Repair (APR) struggles with complex logic errors and silent failures. Current LLM-based APR methods are mostly static, relying on source code and basic test outputs, which fail to accurately capture complex runtime behaviors and dynamic data dependencies. While incorporating runtime evidence like execution traces exposes concrete state transitions, a single LLM interpreting this in isolation often overfits to specific hypotheses, producing patches that satisfy tests by coincidence rather than correct logic. Therefore, runtime evidence should act as objective constraints rather than mere additional input. We propose TraceRepair, a multi-agent framework that leverages runtime facts as shared constraints for patch validation. A probe agent captures execution snapshots of critical variables to form an objective repair basis. Meanwhile, a committee of specialized agents cross-verifies candidate patches to expose inconsistencies and iteratively refine them. Evaluated on the Defects4J benchmark, TraceRepair correctly fixes 392 defects, substantially outperforming existing LLM-based approaches. Extensive experiments demonstrate improved efficiency and strong generalization on a newly constructed dataset of recent bugs, confirming that performance gains arise from dynamic reasoning rather than memorization.
Abstract（参考訳）: 自動プログラム修復(APR)は複雑なロジックエラーとサイレント障害に悩まされる。現在のLLMベースのAPRメソッドはほとんどが静的で、ソースコードと基本的なテスト出力に依存しており、複雑なランタイムの振る舞いや動的データ依存関係を正確にキャプチャできない。実行トレースのような実行時エビデンスを組み込むことは、具体的な状態遷移を露呈するが、これを独立して解釈する単一のLCMは、しばしば特定の仮説に過度に適合し、正しい論理ではなく偶然にテストを満たすパッチを生成する。したがって、実行時のエビデンスは、単なる追加入力ではなく、客観的な制約として振る舞うべきです。我々は、パッチ検証のための共有制約としてランタイム事実を活用するマルチエージェントフレームワークであるTraceRepairを提案する。プローブエージェントは、クリティカル変数の実行スナップショットをキャプチャして、客観的な修復ベースを形成する。一方、専門エージェントの委員会は、不整合を暴露し、それらを反復的に洗練するために、候補パッチを横断的に検証する。 Defects4Jベンチマークで評価すると、TraceRepairは392の欠陥を正しく修正し、既存のLCMベースのアプローチを大幅に上回っている。大規模な実験は、新しく構築された最近のバグのデータセット上で効率の向上と強力な一般化を示し、記憶よりも動的な推論によってパフォーマンスが向上することを確認する。

論文の概要: Runtime Execution Traces Guided Automated Program Repair with Multi-Agent Debate

関連論文リスト