Fugu-MT 論文翻訳(概要): Automated Extract Method Refactoring with Open-Source LLMs: A Comparative Study

論文の概要: Automated Extract Method Refactoring with Open-Source LLMs: A Comparative Study

arxiv url: http://arxiv.org/abs/2510.26480v1
Date: Thu, 30 Oct 2025 13:34:41 GMT
ステータス: 翻訳完了
システム内更新日: 2025-10-31 16:05:09.831749
Title: Automated Extract Method Refactoring with Open-Source LLMs: A Comparative Study
Title（参考訳）: オープンソースのLCMを用いた自動抽出法の比較研究
Authors: Sivajeet Chand, Melih Kilic, Roland Würsching, Sushant Kumar Pandey, Alexander Pretschner,
Abstract要約: 抽出方法(EMR)は、コードの可読性や保守性の改善が重要であるにもかかわらず、依然として困難で手作業がほとんどである。オープンソースのリソース効率の高い大規模言語モデル(LLM)の最近の進歩は、そのようなハイレベルなタスクに対して、有望な新しいアプローチを提供する。
参考スコア（独自算出の注目度）: 35.50372545468027
License: http://creativecommons.org/licenses/by/4.0/
Abstract: Automating the Extract Method refactoring (EMR) remains challenging and largely manual despite its importance in improving code readability and maintainability. Recent advances in open-source, resource-efficient Large Language Models (LLMs) offer promising new approaches for automating such high-level tasks. In this work, we critically evaluate five state-of-the-art open-source LLMs, spanning 3B to 8B parameter sizes, on the EMR task for Python code. We systematically assess functional correctness and code quality using automated metrics and investigate the impact of prompting strategies by comparing one-shot prompting to a Recursive criticism and improvement (RCI) approach. RCI-based prompting consistently outperforms one-shot prompting in test pass rates and refactoring quality. The best-performing models, Deepseek-Coder-RCI and Qwen2.5-Coder-RCI, achieve test pass percentage (TPP) scores of 0.829 and 0.808, while reducing lines of code (LOC) per method from 12.103 to 6.192 and 5.577, and cyclomatic complexity (CC) from 4.602 to 3.453 and 3.294, respectively. A developer survey on RCI-generated refactorings shows over 70% acceptance, with Qwen2.5-Coder rated highest across all evaluation criteria. In contrast, the original code scored below neutral, particularly in readability and maintainability, underscoring the benefits of automated refactoring guided by quality prompts. While traditional metrics like CC and LOC provide useful signals, they often diverge from human judgments, emphasizing the need for human-in-the-loop evaluation. Our open-source benchmark offers a foundation for future research on automated refactoring with LLMs.
Abstract（参考訳）: コードの可読性と保守性を改善することの重要性にもかかわらず、抽出メソッドリファクタリング(EMR)の自動化は依然として難しく、手作業で行われている。オープンソースのリソース効率の高い大規模言語モデル(LLM)の最近の進歩は、そのようなハイレベルなタスクを自動化するための有望な新しいアプローチを提供する。本研究では,Python コードの EMR タスクにおいて,3B から 8B のパラメータサイズにまたがる,最先端のオープンソース LLM を5 つ評価する。自動メトリクスを用いて機能的正当性とコード品質を体系的に評価し,一発のプロンプトを再帰的批判・改善(RCI)アプローチと比較することにより,戦略の促進効果を検討する。 RCIベースのプロンプトは、テストパス率とリファクタリング品質において、ワンショットのプロンプトよりも一貫して優れています。最高のパフォーマンスモデルであるDeepseek-Coder-RCIとQwen2.5-Coder-RCIは、テストパスパーセンテージ(TPP)スコアが0.829と0.808であり、メソッド毎のコード(LOC)は12.103から6.192と5.577、サイクロマティック複雑性(CC)は4.602から3.453と3.294である。 RCI生成リファクタリングに関する開発者調査では70%以上が受け入れており、Qwen2.5-Coderはすべての評価基準の中で最も高い評価を受けている。対照的に、オリジナルのコードは、特に可読性と保守性において中立以下にスコアされ、品質上のプロンプトによって導かれる自動リファクタリングの利点が強調された。 CCやLOCのような従来のメトリクスは有用な信号を提供するが、それらはしばしば人間の判断から分岐し、人間のループ評価の必要性を強調している。私たちのオープンソースベンチマークは、LLMによる自動リファクタリングに関する将来の研究の基盤を提供します。

論文の概要: Automated Extract Method Refactoring with Open-Source LLMs: A Comparative Study

関連論文リスト