Assessing inter-rater agreement among large language model judges: Evidence from persona evaluation across varying educational risk contexts
| dc.contributor.author | Amin, Danial | |
| dc.contributor.author | Salminen, Joni | |
| dc.contributor.author | Jansen, Bernard J. | |
| dc.contributor.department | fi=Ei alustaa|en=No platform| | |
| dc.contributor.orcid | https://orcid.org/0009-0000-7597-2267 | |
| dc.contributor.orcid | https://orcid.org/0000-0003-3230-0561 | |
| dc.date.accessioned | 2026-10-02T12:09:00Z | |
| dc.date.issued | 2026 | |
| dc.description.abstract | Large language models (LLMs) are increasingly proposed as automated judges for evaluation, yet their reliability remains unclear. This study examines the inter-rater agreement of three LLM judges (GPT-4o, Claude 3 Sonnet, and Gemini 1.5 Pro) when evaluating fairness, lack of stereotypicality, and diversity in user personas across varying educational risk contexts. We analyzed 60 individual personas and 12 persona sets to assess whether these judges demonstrate sufficient consistency to function as reliable evaluators. Although inter-rater agreement was moderate at the individual persona level, three critical limitations emerged. At the set level, agreement ranged from .39 to .63 across persona generators and therefore did not show a uniform decline relative to individual-level agreement. Mean dimension-support scores were 41% to 61% lower, although the individual- and set-level tasks are not directly comparable. Second, non-compliance rates, where judges could not provide textual justification for their ratings, ranged from 9.6 to 26.7% at the individual level and increased to 34% to 49% at the set level. Third, agreement differed across the three educational scenarios classified by risk: individual-level agreement declined from the lower- to higher-risk scenarios, while set-level agreement showed no monotonic pattern. We conclude that, under the tested protocol, current LLM judges demonstrate poor-to-moderate inter-rater agreement for persona evaluation. These findings contribute a descriptive characterization of LLM evaluation behavior across risk contexts and provide a foundation for future validation studies comparing LLM agreement against human judgment. | en |
| dc.description.reviewstatus | fi=vertaisarvioitu|en=peerReviewed| | |
| dc.identifier.citation | Amin, D., Salminen, J., & Jansen, B. J. (2027). Assessing inter-rater agreement among large language model judges: Evidence from persona evaluation across varying educational risk contexts. Technological forecasting and social change, 234, Article 124894. https://doi.org/10.1016/j.techfore.2026.124894 | |
| dc.identifier.uri | https://osuva.uwasa.fi/handle/11111/21347 | |
| dc.identifier.urn | URN:NBN:fi-fe20261002130565 | |
| dc.language.iso | en | |
| dc.publisher | Elsevier | |
| dc.relation.doi | https://doi.org/10.1016/j.techfore.2026.124894 | |
| dc.relation.ispartofjournal | Technological forecasting and social change | |
| dc.relation.issn | 1873-5509 | |
| dc.relation.issn | 0040-1625 | |
| dc.relation.url | https://doi.org/10.1016/j.techfore.2026.124894 | |
| dc.relation.url | https://urn.fi/URN:NBN:fi-fe20261002130565 | |
| dc.relation.volume | 234 | |
| dc.rights | https://creativecommons.org/licenses/by/4.0/ | |
| dc.rights.copyright | © 2026 The Authors. Published by Elsevier Inc. This is an open access article under the CC BY license ( http://creativecommons.org/licenses/by/4.0/ ). | |
| dc.source.identifier | 643e7697-7dae-43b1-a725-ef92d7d4cddb | |
| dc.source.metadata | SoleCRIS | |
| dc.subject | Evaluation | |
| dc.subject | LLM as a judge | |
| dc.subject | Large language models | |
| dc.subject | Fairness | |
| dc.subject | Diversity | |
| dc.subject | Lack of stereotypicality | |
| dc.subject.discipline | fi=Markkinointi|en=Marketing| | |
| dc.title | Assessing inter-rater agreement among large language model judges: Evidence from persona evaluation across varying educational risk contexts | |
| dc.type.okm | fi=A1 Alkuperäisartikkeli tieteellisessä aikakauslehdessä (vertaisarvioitu)|en=A1 Journal article (peer-reviewed)| | |
| dc.type.publication | article | |
| dc.type.version | publishedVersion |
Tiedostot
1 - 1 / 1
