The analysis of marking reliability through the approach of gauge repeatability and reproducibility (GR&R) study: a case of English-speaking test

Autor:	Pornphan Sureeyatanapas, Panitas Sureeyatanapas, Uthumporn Panitanarak, Jittima Kraisriwattana, Patchanan Sarootyanapat, Daniel O’Connell
Jazyk:	angličtina
Rok vydání:	2024
Předmět:	GR&R ICC Inter-rater reliability Intra-rater reliability Scoring consistency Speaking proficiency Language and Literature
Zdroj:	Language Testing in Asia, Vol 14, Iss 1, Pp 1-28 (2024)
Druh dokumentu:	article
ISSN:	2229-0443
DOI:	10.1186/s40468-023-00271-z
Popis:	Abstract Ensuring consistent and reliable scoring is paramount in education, especially in performance-based assessments. This study delves into the critical issue of marking consistency, focusing on speaking proficiency tests in English language learning, which often face greater reliability challenges. While existing literature has explored various methods for assessing marking reliability, this study is the first of its kind to introduce an alternative statistical tool, namely the gauge repeatability and reproducibility (GR&R) approach, to the educational context. The study encompasses both intra- and inter-rater reliabilities, with additional validation using the intraclass correlation coefficient (ICC). Using a case study approach involving three examiners evaluating 30 recordings of a speaking proficiency test, the GR&R method demonstrates its effectiveness in detecting reliability issues over the ICC approach. Furthermore, this research identifies key factors influencing scoring inconsistencies, including group performance estimation, work presentation order, rubric complexity and clarity, the student’s chosen topic, accent familiarity, and recording quality. Importantly, it not only pinpoints these root causes but also suggests practical solutions, thereby enhancing the precision of the measurement system. The GR&R method can offer significant contributions to stakeholders in language proficiency assessment, including educational institutions, test developers and policymakers. It is also applicable to other cases of performance-based assessments. By addressing reliability issues, this study provides insights to enhance the fairness and accuracy of subjective judgements, ultimately benefiting overall performance comparisons and decision making.
Databáze:	Directory of Open Access Journals
Externí odkaz:	https://doaj.org/article/3f5d4bbd9b32473ebe7a41a97d55383f Zobrazit plný text záznamu Full text from SpringerLink View record in DOAJ