- Research Article
1
- 10.1108/aci-07-2025-0308
Deep learning approaches for predicting cheating from student exam results: a comparative study under imbalanced data conditions
- Jan 09, 2026
- Applied Computing and Informatics
- Dao Phuc Minh Huy + 2 more +2
Purpose To detect academic misconduct in students' assessment score trajectories under severe class imbalance. The paper compares tabular learners, gradient-boosting models, and sequence-aware deep networks, and proposes a precision–recall–centric evaluation and deployment protocol (calibration, threshold selection and Recall@Top-k%) tailored to rare-event screening in educational settings. Design/methodology/approach A cohort of 1,527 students (2021–2024) is modeled using ten algorithms: LR, DT, RF, SVM, MLP, XGBoost, CatBoost, LightGBM, GRU-RNN and 1D-CNN. Features encode sequential score dynamics and metadata. Models are tuned via cross-validation; probabilities are calibrated (Platt/Isotonic); operating thresholds are chosen on validation to maximize minority-class F1 or a cost-sensitive utility. Performance is assessed on a hold-out test set with PR-AUC (headline), F1(+), Recall@Top-k%, ROC-AUC, calibration curves, Brier, and bootstrap CIs. Findings Sequence-aware models dominate: GRU-RNN and 1D-CNN achieve ROC-AUC ˜0.97–0.98 and the highest F1(+) and Recall@Top-5%. Tabular/boosting baselines show ˜0.90 accuracy yet miss most positives at the default 0.5 threshold, highlighting the necessity of calibration and threshold optimization. With PR-centric selection and tuned operating points, deep temporal models yield strong screening utility for limited human review budgets. Research limitations/implications Labels reflect suspected–not adjudicated–cheating, introducing noise. The single-institution cohort may limit external validity; temporal shift across semesters can degrade performance. Future work should include multi-site evaluation, collusion/graph modeling, semi-/weak-supervision for noisy labels and governance topics (fairness audits, drift monitoring and uncertainty reporting). Practical implications This study highlights how educational institutions can leverage machine learning for early detection of academic dishonesty based on historical performance data. CNN and RNN models are promising tools for identifying anomalous learning patterns. However, practical deployment requires preprocessing techniques to manage class imbalance and threshold optimization to reduce false negatives. The findings provide a roadmap for building automated cheating detection systems in both online and traditional assessment environments. Social implications By improving the ability to detect cheating, this research contributes to fairer academic environments, upholding educational integrity and credibility of credentials. However, ethical considerations must be taken into account to avoid false accusations and ensure student rights. Human-in-the-loop systems are crucial for verifying algorithmic predictions before disciplinary action, thereby fostering transparency and accountability in automated decision-making processes. Originality/value The study unifies a minority-focused evaluation protocol with a comprehensive comparison of tabular, boosting, and sequence-aware models for cheating detection from score trajectories. It demonstrates the decisive value of temporal representation learning and provides a reproducible pipeline and operational metrics that align model performance with real investigative workflows.
Read more