- Research Article
- 10.1109/access.2026.3657348
Performance Evaluation of OpenAI’s o4-mini on CSAT Mathematics: Multimodal Reasoning and the Reasoning–Verification–Reanalysis Loop
- Jan 01, 2026
- IEEE Access
- Sejun Oh + 2 more +2
Recent advances in multimodal large language models (LLMs) have expanded their potential for solving complex mathematical problems; however, their reasoning capabilities in real, curriculum-aligned examinations remain underexplored. This study evaluates the mathematical reasoning performance of OpenAI’s o4-mini model on the mathematics section of the Korean College Scholastic Ability Test (KCSAT), using official 2023–2025 exams and mock tests as benchmark data. The o4-mini is a multimodal LLM capable of processing both textual and diagrammatic information, enabling assessment on problems that require integrated visual reasoning. Quantitative analyses reveal that o4-mini demonstrates strong performance on standard problems and moderate success on high-difficulty items, while qualitative analyses uncover a distinctive “Reasoning–Verification–Reanalysis (RVR) loop” in its reasoning process. This iterative pattern—where the model reasons, self-verifies intermediate steps, and reanalyzes discrepancies—reflects a human-like metacognitive behavior that often improves problem-solving accuracy. These findings provide new insights into the cognitive dynamics of multimodal LLMs, highlighting both their potential and limitations in high-stakes mathematical assessments, and suggesting pathways for enhancing AI reliability through structured reasoning and self-verification mechanisms.
Read more