Comparing Teacher and Artificial Intelligence Scoring in Writing Assessment: A Generalizability Theory Analysis


Asma B.

LANGUAGE TEACHING RESEARCH, cilt.0, sa.0, ss.1-26, 2026 (SSCI, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 0 Sayı: 0
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1177/13621688261463602
  • Dergi Adı: LANGUAGE TEACHING RESEARCH
  • Derginin Tarandığı İndeksler: Scopus, Social Sciences Citation Index (SSCI)
  • Sayfa Sayıları: ss.1-26
  • Akdeniz Üniversitesi Adresli: Evet

Özet

This study examined the use of artificial intelligence tools, which have garnered significant attention in recent years, in the assessment and evaluation processes of language education. For this purpose, student essays were scored by Turkish middle school teachers and artificial intelligence tools both with and without the use of a rubric, and the findings were evaluated based on generalizability theory. Additionally, the research findings were shared with participants to gather qualitative data, which were analysed using the inductive thematic analysis method to support the research results. The findings revealed that in evaluations conducted without a rubric, teachers were limited in their ability to distinguish individual differences and demonstrated low scoring consistency. In contrast, artificial intelligence tools were more effective in distinguishing individual differences and exhibited high consistency. In evaluations conducted using a rubric, scoring consistency increased in both groups, although, as in the first evaluation, artificial intelligence tools demonstrated a higher level of consistency. Regarding the research findings, teachers expressed that individual biases, mood, and professional experiences influenced their scoring processes and emphasized the potential of rubrics and artificial intelligence-supported feedback systems for achieving more consistent results. Artificial intelligence tools, on the other hand, highlighted their independence from subjective factors but stressed the need for more diverse and generalizable datasets to further enhance their evaluation capacities.