A Large-Scale Empirical Evaluation of LLMs for Automated Self-Admitted Technical Debt Repayment
Self-Admitted Technical Debt (SATD), cases where developers intentionally acknowledge suboptimal solutions in code through comments, poses a significant challenge to software maintainability. Left unresolved, SATD can degrade code quality, increase maintenance costs, and hinder long-term software sustainability. Manually identifying and resolving SATD is time-consuming and error-prone, often competing with feature development and release deadlines. Automating SATD repayment is therefore essential to alleviate developer burden. While Large Language Models (LLMs) have shown promise in tasks like code generation and program repair, their potential in automated SATD repayment remains underexplored. In this paper, we identify three key technical challenges in training and evaluating LLMs for SATD repayment: (1) dataset representativeness and scalability, (2) removal of irrelevant SATD repayment samples, and (3) limitations of existing evaluation metrics, i.e., BLEU and CrystalBLEU. To address the first two dataset-related challenges, we adopt a language-independent SATD tracing tool and design a 10-step filtering pipeline to extract SATD repayments from repository commit histories, leveraging LLM-as-judge for relevance filtering. This results in two new large-scale SATD repayment datasets: 58,722 items for Python and 97,347 items for Java. To improve evaluation, we introduce two diff-based metrics, BLEU-diff and CrystalBLEU-diff, which measure code changes rather than whole code, thereby mitigating the impact of SATD-irrelevant source code. Additionally, we propose a new metric, Line-Level Exact Match on Diff (LEMOD), which is both interpretable and informative, providing fine-grained insights into SATD repayment quality. Using our new benchmarks and evaluation metrics, we evaluate two types of automated SATD repayment methods: fine-tuning smaller models (a few hundred million parameters) and prompt engineering with large-scale models, including four SOTA open-source LLMs and GPT-4o-mini. Our results reveal that while fine-tuned smaller models achieve Exact Match (EM) scores comparable to prompt-based approaches, they underperform on BLEU-based metrics and LEMOD. The best EM performance is achieved by Gemma-2-9B, correctly addressing 10.1% of Python SATDs and 8.1% of Java SATDs with a simple prompt. The only previous study in this domain focused solely on Java, achieving an EM score of 2.3%. When using BLEU-diff, CrystalBLEU-diff, and LEMOD metrics, Llama-3.1-70B-Instruct and GPT-4o-mini deliver the highest performance. While our three proposed metrics demonstrate strong correlations with EM (0.65~0.84), both BLEU and CrystalBLEU show almost no correlation with this intuitive metric (0.01~0.08), highlighting their largely unpredictable behavior when applied directly to the entire code rather than the diff. Our work contributes a robust benchmark, improved evaluation metrics, and a comprehensive evaluation of LLMs, advancing research on automated SATD repayment.
Read more