- Research Article
- 10.1080/1750399x.2026.2660553
ChatGPT as an evaluator in summative assessment of interpreting: a psychometric analysis of ChatGPT versus human raters
- Apr 22, 2026
- The Interpreter and Translator Trainer
- Xiaolei Lu + 1 more +1
ABSTRACT Translation and interpreting (T&I) have long been taught and assessed in higher education contexts worldwide. With the growing demand for grading student-generated T&I, researchers are focusing on automatic T&I assessment. In this article, we aim to initially investigate the utility of ChatGPT as an automatic evaluator for interpreting, by comparing ChatGPT-based scores with those of human raters. We recruited 45 raters from diverse backgrounds to assess 56 English-Chinese student interpretations. Twenty-four ChatGPT-based e-raters were also configured to assess the same interpretations, varying the reference rendition availability, scoring granularity, and model randomness. We conducted statisticalanalysis to examine the reliability, severity, validity, and accuracy of ChatGPT-based e-raters. Our findings reveal that: a) e-raters showed greater reliability and severity than human raters; b) e-raters generally matched human raters in concurrent validity; and c) e-raters outperformed some of their human counterparts in scoring accuracy for English-to-Chinese interpretations. Among the 24 e-raters, those scoring at the passage level exhibited lower reliability, severity and validity but, interestingly, achieved more accurate scores for English-to-Chinese interpretations. These results suggest that ChatGPT-based e-raters may be used as a secondary rater for cross-validation alongside human raters but should not be deployed as sole scorers.
Read more