Conference paper 2021

A Multidimensional Item Response Theory Model for Rubric-Based Writing Assessment

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Conference · Vol. 12748 LNAI · pp. 420-432
Abstract

When human raters grade student writing assignments, writing assessment often involves the use of a scoring rubric consisting of multiple evaluation items in order to increase the objectivity of evaluation. However, even when using a rubric, assigned scores are known to be influenced by the characteristics of both the rubric’s evaluation items and the raters, thus decreasing the reliability of student assessment. To resolve this problem, these characteristic effects have been considered in many recently proposed item response theory (IRT) models for estimating student ability. Such IRT models assume unidimensionality, meaning that a rubric measures one latent ability; in practice, however, this assumption might not be satisfied because a rubric’s evaluation items are often designed to measure multiple sub-abilities that constitute a targeted ability. To address this issue, this study proposes a multidimensional extension of such an IRT model for rubric-based writing assessment. The proposed model improves the assessment reliability. Furthermore, the model is useful for objective and detailed analysis of rubric quality and its construct validity. This study demonstrates the effectiveness of the proposed model through simulation experiments and application to real data. © 2021, Springer Nature Switzerland AG.

Keywords

Author Keywords

Educational/psychological measurement Analytic rubrics Statistical/probabilistic model Test theory Writing assessment

Index Keywords

Quality control Students Item response theory Analytic rubric Educational/psychological measurement Evaluation items Probabilistic models Psychological measurements Statistical/probabilistic model Test theories Theory model Writing assessment
Author Affiliations
The University of Electro-Communications, Chofu, Tokyo, Japan
Funding & Acknowledgements
No funding information
References 10 References
1 International Journal of Secondary Education, (2016)
2 Bernardin, H. John, Rater Rating-Level Bias and Accuracy in Performance Appraisals: The Impact OF Rater Personality, Performance Management Competence, and Rater Accountability, Human Resource Management, 55, 2, pp. 321-340, (2016)
3 Brooks, Steve, Handbook of Markov Chain Monte Carlo, Handbook of Markov Chain Monte Carlo, pp. 1-592, (2011)
4 Carpenter, Bob, Stan: A probabilistic programming language, Journal of Statistical Software, 76, 1, (2017)
5 Ielts Research Reports Online Series, (2017)
6 Deng, Sien, Extreme Response Style and the Measurement of Intra-Individual Variability in Affect, Multivariate Behavioral Research, 53, 2, pp. 199-218, (2018)
7 Eckes, Thomas, Introduction to many-facet rasch measurement: Analyzing and evaluating rater-mediated assessments: Second edition, Introduction to Many-Facet Rasch Measurement: Analyzing and Evaluating Rater-Mediated Assessments: Second Edition, 22, pp. 1-241, (2015)
8 Bayesian Item Response Modeling Theory and Applications, (2010)
9 Gelman, Andrew E., Bayesian data analysis, third edition, Bayesian Data Analysis, Third Edition, pp. 1-646, (2013)
10 Gelman, Andrew E., Inference from iterative simulation using multiple sequences, Statistical Science, 7, 4, pp. 457-472, (1992)
Quick Actions
Full Text via DOI
Citation Metrics
0
Times Cited (Scopus)

References 10
Document Identifiers