• Home
  • Search
  • Demystifying LLM-Based Software Engineering Agents
  • Cite Icon35
  • https://doi.org/10.1145/3715754Copy DOI Icon

Demystifying LLM-Based Software Engineering Agents

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

Recent advancements in large language models (LLMs) have significantly advanced the automation of software development tasks, including code synthesis, program repair, and test generation. More recently, researchers and industry practitioners have developed various autonomous LLM agents to perform end-to-end software development tasks. These agents are equipped with the ability to use tools, run commands, observe feedback from the environment, and plan for future actions. However, the complexity of these agent-based approaches, together with the limited abilities of current LLMs, raises the following question: Do we really have to employ complex autonomous software agents? To attempt to answer this question, we build Agentless – an agentless approach to automatically resolve software development issues. Compared to the verbose and complex setup of agent-based approaches, Agentless employs a simplistic three-phase process of localization, repair, and patch validation, without letting the LLM decide future actions or operate with complex tools. Our results on the popular SWE-bench Lite benchmark show that surprisingly the simplistic Agentless is able to achieve both the highest performance (32.00%, 96 correct fixes) and low cost ($0.70) compared with all existing open-source software agents at the time of paper submission! Agentless also achieves more than 50% solve rate when using Claude 3.5 Sonnet on the new SWE-bench Verified benchmark. In fact, Agentless has already been adopted by OpenAI as the go-to approach to showcase the real-world coding performance of both GPT-4o and the new o1 models; more recently, Agentless has also been used by DeepSeek to evaluate their newest DeepSeek V3 and R1 models. Furthermore, we manually classified the problems in SWE-bench Lite and found problems with exact ground truth patches or insufficient/misleading issue descriptions. As such, we construct SWE-bench Lite-𝑆 by excluding such problematic issues to perform more rigorous evaluation and comparison. Our work highlights the currently overlooked potential of a simplistic, cost-effective technique in autonomous software development. We hope Agentless will help reset the baseline, starting point, and horizon for autonomous software agents, and inspire future work along this crucial direction. We have open-sourced Agentless at: https://github.com/OpenAutoCoder/Agentless

Similar Papers
  • Research Article

LLM-Based Web Generation Quality Assessment

  • Mar 27, 2025
  • Theoretical and Natural Science
  • Yizhen Gong +1
  • Research Article

Benchmarking LLMs for Unit Test Generation from Real-World Functions

  • Mar 28, 2026
  • ACM Transactions on Software Engineering and Methodology
  • Dong Huang +5
  • Conference Article

LORE: Continual Logit Rewriting Fosters Faithful Generation

  • Jan 01, 2025
  • Charles Yu +4
  • Research Article
  • Citations3

Less Is More: On the Importance of Data Quality for Unit Test Generation

  • Jun 19, 2025
  • Proceedings of the ACM on Software Engineering
  • Junwei Zhang +5
  • Research Article
  • Citations1

AI-Powered Code Generation Evaluating the Effectiveness of Large Language Models (LLMs) in Automated Software Development

  • Mar 31, 2023
  • Journal of Artificial Intelligence & Cloud Computing
  • Ravikanth Konda
  • Research Article
  • Citations14

Improving Knowledge Extraction from LLMs for Task Learning through Agent Analysis

  • Mar 24, 2024
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • James R Kirk +3
  • Research Article
  • Citations1

Limitations and mitigation strategies for using generative artificial intelligence in medical writing: a narrative review

  • Feb 10, 2026
  • Journal of Korean Medical Association
  • Ki-Hyun Jeon
  • Conference Article
  • Citations8

Evaluating Fault Localization and Program Repair Capabilities of Existing Closed-Source General-Purpose LLMs

  • Apr 20, 2024
  • Shengbei Jiang +5
  • Research Article
  • Citations17

ContrastRepair : Enhancing Conversation-Based Automated Program Repair via Contrastive Test Case Pairs

  • Oct 03, 2025
  • ACM Transactions on Software Engineering and Methodology
  • Jiaolong Kong +5
  • Research Article
  • Citations61

Large language models illuminate a progressive pathway to artificial intelligent healthcare assistant

  • May 17, 2024
  • Medicine Plus
  • Mingze Yuan +11
  • Research Article
  • Citations11

Early Formalization of AI-tools Usage in Software Engineering in Europe: Study of 2023

  • Dec 08, 2023
  • International Journal of Information Technology and Computer Science
  • Denis S Pashchenko
  • Conference Article
  • Citations4

NegativePrompt: Leveraging Psychology for Large Language Models Enhancement via Negative Emotional Stimuli

  • Aug 01, 2024
  • Xu Wang +2
  • Research Article

A Large-Scale Empirical Evaluation of LLMs for Automated Self-Admitted Technical Debt Repayment

  • Feb 10, 2026
  • ACM Transactions on Software Engineering and Methodology
  • Mohammad Sadegh Sheikhaei +3
  • Research Article

Securing social network user data in large language model deployments: challenges and best practices

  • Sep 11, 2025
  • Cluster Computing
  • Nasir Ahmad Jalali +1
  • Research Article

Analyzing LLM-generated code according to four ISO/IEC 5055:2021 categories

  • Jan 01, 2025
  • IEEE Access
  • Rasmus Krebs +1
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.