Assessing inter-rater agreement among large language model judges: Evidence from persona evaluation across varying educational risk contexts

dc.contributor.authorAmin, Danial
dc.contributor.authorSalminen, Joni
dc.contributor.authorJansen, Bernard J.
dc.contributor.departmentfi=Ei alustaa|en=No platform|
dc.contributor.orcidhttps://orcid.org/0009-0000-7597-2267
dc.contributor.orcidhttps://orcid.org/0000-0003-3230-0561
dc.date.accessioned2026-10-02T12:09:00Z
dc.date.issued2026
dc.description.abstractLarge language models (LLMs) are increasingly proposed as automated judges for evaluation, yet their reliability remains unclear. This study examines the inter-rater agreement of three LLM judges (GPT-4o, Claude 3 Sonnet, and Gemini 1.5 Pro) when evaluating fairness, lack of stereotypicality, and diversity in user personas across varying educational risk contexts. We analyzed 60 individual personas and 12 persona sets to assess whether these judges demonstrate sufficient consistency to function as reliable evaluators. Although inter-rater agreement was moderate at the individual persona level, three critical limitations emerged. At the set level, agreement ranged from .39 to .63 across persona generators and therefore did not show a uniform decline relative to individual-level agreement. Mean dimension-support scores were 41% to 61% lower, although the individual- and set-level tasks are not directly comparable. Second, non-compliance rates, where judges could not provide textual justification for their ratings, ranged from 9.6 to 26.7% at the individual level and increased to 34% to 49% at the set level. Third, agreement differed across the three educational scenarios classified by risk: individual-level agreement declined from the lower- to higher-risk scenarios, while set-level agreement showed no monotonic pattern. We conclude that, under the tested protocol, current LLM judges demonstrate poor-to-moderate inter-rater agreement for persona evaluation. These findings contribute a descriptive characterization of LLM evaluation behavior across risk contexts and provide a foundation for future validation studies comparing LLM agreement against human judgment.en
dc.description.reviewstatusfi=vertaisarvioitu|en=peerReviewed|
dc.identifier.citationAmin, D., Salminen, J., & Jansen, B. J. (2027). Assessing inter-rater agreement among large language model judges: Evidence from persona evaluation across varying educational risk contexts. Technological forecasting and social change, 234, Article 124894. https://doi.org/10.1016/j.techfore.2026.124894
dc.identifier.urihttps://osuva.uwasa.fi/handle/11111/21347
dc.identifier.urnURN:NBN:fi-fe20261002130565
dc.language.isoen
dc.publisherElsevier
dc.relation.doihttps://doi.org/10.1016/j.techfore.2026.124894
dc.relation.ispartofjournalTechnological forecasting and social change
dc.relation.issn1873-5509
dc.relation.issn0040-1625
dc.relation.urlhttps://doi.org/10.1016/j.techfore.2026.124894
dc.relation.urlhttps://urn.fi/URN:NBN:fi-fe20261002130565
dc.relation.volume234
dc.rightshttps://creativecommons.org/licenses/by/4.0/
dc.rights.copyright© 2026 The Authors. Published by Elsevier Inc. This is an open access article under the CC BY license ( http://creativecommons.org/licenses/by/4.0/ ).
dc.source.identifier643e7697-7dae-43b1-a725-ef92d7d4cddb
dc.source.metadataSoleCRIS
dc.subjectEvaluation
dc.subjectLLM as a judge
dc.subjectLarge language models
dc.subjectFairness
dc.subjectDiversity
dc.subjectLack of stereotypicality
dc.subject.disciplinefi=Markkinointi|en=Marketing|
dc.titleAssessing inter-rater agreement among large language model judges: Evidence from persona evaluation across varying educational risk contexts
dc.type.okmfi=A1 Alkuperäisartikkeli tieteellisessä aikakauslehdessä (vertaisarvioitu)|en=A1 Journal article (peer-reviewed)|
dc.type.publicationarticle
dc.type.versionpublishedVersion

Tiedostot

Näytetään 1 - 1 / 1
Ladataan...
Name:
nbnfi-fe20261002130565.pdf
Size:
907.68 KB
Format:
Adobe Portable Document Format

Kokoelmat