“Half My Survey Data Is Bad; The Problem Is, I Can't Tell Which Half”: Evaluating the Validatability of Large-Scale Survey Data Using 176 Algorithmically Generated Personas
| dc.contributor.author | Jansen, Bernard J. | |
| dc.contributor.author | Akari, Marwan | |
| dc.contributor.author | Amin, Safa | |
| dc.contributor.author | Jung, Soon-Gyo | |
| dc.contributor.author | Amin, Danial | |
| dc.contributor.author | Salminen, Joni | |
| dc.contributor.editor | Brooker, Sam | |
| dc.contributor.editor | Benatti, Francesca | |
| dc.contributor.editor | Pisarski, Mariusz | |
| dc.contributor.editor | Adamou, Alessandro | |
| dc.contributor.orcid | https://orcid.org/0009-0000-7597-2267 | |
| dc.contributor.orcid | https://orcid.org/0000-0003-3230-0561 | |
| dc.date.accessioned | 2026-09-23T08:58:00Z | |
| dc.date.issued | 2026 | |
| dc.description.abstract | Survey data is foundational to much user research, including design artifacts such as personas. However, if survey data is invalidatable, the credibility of any downstream analysis is fundamentally undermined. This work starts with the premise that distinguishing valid from invalid survey data solely through internal survey checks, such as attention checks, is practically infeasible. We conducted a large-scale empirical survey of social media users (N ≈ 20,000) and used persona creation as an analytical lens. Of these responses, 8,140 were classified as (ostensibly) valid and 11,860 as invalid based on passing or failing attention checks. We construct ten datasets by progressively replacing ‘valid’ responses with ‘invalid’ ones in increments of 10%. From each dataset, we generate 16 personas, resulting in a total of 176, and compare their divergence. Results were stable despite data degradation. Persona demographics remained unchanged in most conditions, with demographic consistency at nearly 90%; an average of 12% of persona attributes were unchanged. More than 80% of the survey item diversity was consistent across datasets. The findings are that large-scale survey data may be unverifiable through internal checks alone due to the Validity Masking Effect, and that artifacts in such data can obscure underlying data quality issues, highlighting the risks of relying solely on survey data. | en |
| dc.description.reviewstatus | fi=vertaisarvioitu|en=peerReviewed| | |
| dc.format.pagerange | 368-377 | |
| dc.identifier.citation | Jansen, B. J., Akari, M., Amin, S., Jung, S.-G., Amin, D., & Salminen, J. (2026). “Half My Survey Data Is Bad; The Problem Is, I Can't Tell Which Half”: Evaluating the Validatability of Large-Scale Survey Data Using 176 Algorithmically Generated Personas. In S. Brooker, F. Benatti, M. Pisarski, & A. Adamou (Eds.), HT '26: Proceedings of the 37th ACM Conference on Hypertext (pp. 368-377). Association for Computing Machinery. https://doi.org/10.1145/3800935.3830835 | |
| dc.identifier.isbn | 979-8-4007-2564-7 | |
| dc.identifier.uri | https://osuva.uwasa.fi/handle/11111/21325 | |
| dc.identifier.urn | URN:NBN:fi-fe20260923127954 | |
| dc.language.iso | en | |
| dc.publisher | ACM | |
| dc.relation.conference | ACM Conference on Hypertext (HT) | |
| dc.relation.doi | https://doi.org/10.1145/3800935.3830835 | |
| dc.relation.ispartof | HT '26: Proceedings of the 37th ACM Conference on Hypertext | |
| dc.relation.url | https://doi.org/10.1145/3800935.3830835 | |
| dc.relation.url | https://urn.fi/URN:NBN:fi-fe20260923127954 | |
| dc.rights | https://creativecommons.org/licenses/by/4.0/ | |
| dc.rights.copyright | © 2026 Copyright held by the owner/author(s). This work is licensed under a Creative Commons Attribution 4.0 International License. https://creativecommons.org/licenses/by/4.0 | |
| dc.source.identifier | 7e0cd80e-d13d-48f5-b9dc-e62f98f81322 | |
| dc.source.metadata | SoleCRIS | |
| dc.subject | survey data | |
| dc.subject | data validity | |
| dc.subject | automatic persona generation | |
| dc.subject | data quality | |
| dc.subject.discipline | fi=Markkinointi|en=Marketing| | |
| dc.title | “Half My Survey Data Is Bad; The Problem Is, I Can't Tell Which Half”: Evaluating the Validatability of Large-Scale Survey Data Using 176 Algorithmically Generated Personas | |
| dc.type.okm | fi=A4 Vertaisarvioitu artikkeli konferenssijulkaisussa|en=A4 Article in conference proceedings (peer-reviewed)| | |
| dc.type.publication | article | |
| dc.type.version | publishedVersion |
Tiedostot
1 - 1 / 1
