Large language model (LLM)-based auto-graders, like Claude 3.5 Sonnet, show promise in educational technology. To test their capabilities, we conducted an experiment in which four researchers solved five mathematical programming problems in Python, ranging from easy to complex. The researchers received feedback from Claude, using a 22 -criteria rubric designed to evaluate the following key aspects of coding: “Correctness,” “Efficiency,” “Data Structure Usage,” “Code Readability,” and “Testing” before resubmitting their assignments. The mean score improvement of $\mathbf{1 7 . 5}$ points demonstrates that LLMs can improve code quality, with the highest percentage increases in time complexity $(\mathbf{+} \mathbf{2 5 . 4 5 \%})$, efficiency $(\mathbf{+ 2 2 . 5 9 \%})$, and edge case handling ($+22 \%$).