• Home
  • Search
  • Evaluating Language Models for Generating and Judging Programming Feedback
  • Cite Icon14
  • https://doi.org/10.1145/3641554.3701791Copy DOI Icon

Evaluating Language Models for Generating and Judging Programming Feedback

  • Feb 12, 2025
  • Charles Koutcheme +6 more
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

The emergence of large language models (LLMs) has transformed research and practice across a wide range of domains. Within the computing education research (CER) domain, LLMs have garnered significant attention, particularly in the context of learning programming. Much of the work on LLMs in CER, however, has focused on applying and evaluating proprietary models. In this article, we evaluate the efficiency of open-source LLMs in generating high-quality feedback for programming assignments and judging the quality of programming feedback, contrasting the results with proprietary models. Our evaluations on a dataset of students' submissions to introductory Python programming exercises suggest that state-of-the-art open-source LLMs are nearly on par with proprietary models in both generating and assessing programming feedback. Additionally, we demonstrate the efficiency of smaller LLMs in these tasks and highlight the wide range of LLMs accessible, even for free, to educators and practitioners.

Similar Papers
  • Supplementary Content

Large language and vision-language models for robot: safety challenges, mitigation strategies and future directions

  • Jul 29, 2025
  • Industrial Robot: the international journal of robotics research and application
  • Xiangyu Hu +1
  • Research Article

Influence of structured output constraints on GPT-5-Thinking, Gemini 2.5 Pro, and open-weight LLMs for radiology protocol selection.

  • Apr 10, 2026
  • European radiology experimental
  • Mohammed Bahaaeldin +9
  • Research Article

Unlocking the Potential of Large Language Models in Education: Factors Influencing Adoption by Instructional Designers and Academics

  • Jan 01, 2026
  • Journal of Information Technology Education: Research
  • Katherine L Fourie +2
  • Research Article

Evaluating gpt-4 for zero-shot classification of bleeding and clotting events: Can large language models serve as second reviewers?

  • Nov 03, 2025
  • Blood
  • Samantha Rizzo +5
  • Research Article

Research and selection of Large Learning Models for automation of ABAP-code migration

  • Sep 24, 2025
  • Management of Development of Complex Systems
  • Oleg Pozdnyakov +1
  • PDF
  • Conference Article
  • Citations2

Test Large Language Models on Driving Theory Knowledge and Skills for Connected Autonomous Vehicles

  • Nov 18, 2024
  • Zuoyin Tang +5
  • Research Article

Comparison of proprietary and fine-tuned large language models for multi-label classification of billing codes from radiology reports.

  • Mar 14, 2026
  • European radiology
  • Kamyar Arzideh +12
  • Research Article
  • Citations1

Optimization of traditional methods for determining the similarity of project names and purchases using large language models

  • Apr 01, 2024
  • Litera
  • Aleksei Aleksandrovich Golikov +2
  • Research Article

Readability & quality of large language model responses in CAR-T patient education

  • Nov 03, 2025
  • Blood
  • Sridhar Balasubramanian +3
  • Research Article
  • Citations7

Performance and Reproducibility of Large Language Models in Named Entity Recognition: Considerations for the Use in Controlled Environments

  • Dec 11, 2024
  • Drug Safety
  • Jürgen Dietrich +1
  • Conference Article
  • Citations18

Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?

  • Jan 01, 2024
  • Daniel P Jeong +3
  • Research Article
  • Citations42

Improving accuracy of GPT-3/4 results on biomedical data using a retrieval-augmented language model.

  • Aug 21, 2024
  • PLOS digital health
  • David Soong +10
  • Research Article
  • Citations178

The ethics of disclosing the use of artificial intelligence tools in writing scholarly manuscripts

  • Jun 15, 2023
  • Research Ethics
  • Mohammad Hosseini +2
  • Research Article

Application of Large Language Models (LLMs) to Geriatric Practice and Its Evaluation at 4 VA GRECCs

  • Dec 01, 2025
  • Innovation in Aging
  • Huai Cheng +2
  • Research Article

1401 Bringing Medicine Expertise to Your Screen: A New Frontier in Curbside Sleep Consultation Leveraging Large Language Models?

  • May 19, 2025
  • SLEEP
  • Nina Kuei +4
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.