Assessing inter-rater agreement among large language model judges: Evidence from persona evaluation across varying educational risk contexts
Lopullinen julkaistu versio - 907.68 KB
Amin, D., Salminen, J., & Jansen, B. J. (2027). Assessing inter-rater agreement among large language model judges: Evidence from persona evaluation across varying educational risk contexts. Technological forecasting and social change, 234, Article 124894. https://doi.org/10.1016/j.techfore.2026.124894
© 2026 The Authors. Published by Elsevier Inc. This is an open access article under the CC BY license ( http://creativecommons.org/licenses/by/4.0/ ).
Pysyvä osoite
Kuvaus
Large language models (LLMs) are increasingly proposed as automated judges for evaluation, yet their reliability remains unclear. This study examines the inter-rater agreement of three LLM judges (GPT-4o, Claude 3 Sonnet, and Gemini 1.5 Pro) when evaluating fairness, lack of stereotypicality, and diversity in user personas across varying educational risk contexts. We analyzed 60 individual personas and 12 persona sets to assess whether these judges demonstrate sufficient consistency to function as reliable evaluators. Although inter-rater agreement was moderate at the individual persona level, three critical limitations emerged. At the set level, agreement ranged from .39 to .63 across persona generators and therefore did not show a uniform decline relative to individual-level agreement. Mean dimension-support scores were 41% to 61% lower, although the individual- and set-level tasks are not directly comparable. Second, non-compliance rates, where judges could not provide textual justification for their ratings, ranged from 9.6 to 26.7% at the individual level and increased to 34% to 49% at the set level. Third, agreement differed across the three educational scenarios classified by risk: individual-level agreement declined from the lower- to higher-risk scenarios, while set-level agreement showed no monotonic pattern. We conclude that, under the tested protocol, current LLM judges demonstrate poor-to-moderate inter-rater agreement for persona evaluation. These findings contribute a descriptive characterization of LLM evaluation behavior across risk contexts and provide a foundation for future validation studies comparing LLM agreement against human judgment.
Emojulkaisu
ISBN
ISSN
1873-5509
0040-1625
0040-1625
Aihealue
Kausijulkaisu
Technological forecasting and social change|234
OKM-julkaisutyyppi
A1 Alkuperäisartikkeli tieteellisessä aikakauslehdessä (vertaisarvioitu)
