A private secondary school in Penang is set to introduce artificial intelligence (AI) to mark English essays starting in October, prompting debate about the role of AI in educational assessment.
Proponents argue that AI could offer faster feedback for students and allow teachers to dedicate more time to instruction and student support. However, experts caution that grading essays involves more than identifying grammatical errors or verifying the inclusion of certain points; it requires nuanced judgment about clarity, reasoning, originality, and communication effectiveness—areas where human insight remains critical.
Dr. Ong Yew Chuan, deputy dean of the Faculty of Informatics & Computing at Universiti Sultan Zainal Abidin, emphasized the importance of differentiating between AI as an aid to teachers and AI as the sole evaluator. Drawing on conversations with a British university, Dr. Ong noted that some institutions have declined to use AI for grading partly because students expressed discomfort with machine-only evaluation, underscoring that trust is a fundamental component in assessment.
Before fully implementing AI-based grading, Dr. Ong recommends schools conduct pilot programs where both teachers and AI independently assess the same essays. Such trials could identify areas of agreement and divergence, determine how easily teachers can contest AI evaluations, and assess whether AI enhances the quality of feedback or simply accelerates marking.
Further questions merit consideration: Should students be informed when AI contributes to their grades? Can AI determine important grades without substantial human oversight? In cases of conflicting judgments between teacher and AI, which should take precedence? Additionally, accountability arises when AI systems err. Transparency around the use of AI in grading, akin to requirements for students disclosing AI use in assignments, is another point for discussion.
While there is potential for AI to become a valuable tool in assessment, education leaders stress that its adoption must be grounded in evidence, transparency, and human accountability. The broader query has shifted from whether AI is capable of grading essays to whether its use ultimately improves educational outcomes.
