When Cohen’s Kappa Is Not Enough: Exploring Methods for Estimating Inter-Rater Reliability for Time Sequential Data
Page: 286
We identify a current pragmatic and methodological problem facing RUME researchers who study processes that unfold over time, like classrooms and interviews, with qualitative coding techniques. In many cases, existing methods for estimating agreement between raters fail to account for the properties of this kind of data, are not compatible with how the coding schemes are applied, and fail to properly estimate rater agreement. Despite many peer-reviewers requesting estimates of agreement for coded data, there is often no suitable value to report and researchers must make do with claims about coming to consensus. We review methods found in the literature and evaluate their suitability, strengths, and weaknesses. While each reviewed method is appropriate for some aspect of this kind of data, none satisfies all desired criteria.