‘This study examines the consistency and agreement between generative AI (GenAI)-enabled and human assessments of written assignments in a teacher development course on service-learning. Using a rubric with nine elements, thirty-seven proposals were independently graded by human assessors and OpenAI GPT-4.1. Results show that GenAI tended to assign higher grades than humans, particularly in elements requiring nuanced judgement, such as civic learning and project feasibility. While overall scores demonstrated slight reliability, agreement at the element level was generally poor, with discrepancies most pronounced in subjective criteria. The study also identified a proportional bias, with GenAI grading more generously at lower and mid-range scores. Findings suggest that zero-shot GenAI grading is unlikely to yield fair or reliable results, prompting the need for future research to identify task-specific conditions under which GenAI assessment may be appropriate and effective.’