Collecting Qualitative Data at Scale with Large Language Models: A Case Study

September 18, 2023 · Declared Dead · 🏛 Proc. ACM Hum. Comput. Interact.

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Alejandro Cuevas, Jennifer V. Scurrell, Eva M. Brown, Jason Entenmann, Madeleine I. G. Daepp arXiv ID 2309.10187 Category cs.HC: Human-Computer Interaction Citations 5 Venue Proc. ACM Hum. Comput. Interact. Last Checked 4 months ago

Abstract

Chatbots have shown promise as tools to scale qualitative data collection. Recent advances in Large Language Models (LLMs) could accelerate this process by allowing researchers to easily deploy sophisticated interviewing chatbots. We test this assumption by conducting a large-scale user study (n=399) evaluating 3 different chatbots, two of which are LLM-based and a baseline which employs hard-coded questions. We evaluate the results with respect to participant engagement and experience, established metrics of chatbot quality grounded in theories of effective communication, and a novel scale evaluating "richness" or the extent to which responses capture the complexity and specificity of the social context under study. We find that, while the chatbots were able to elicit high-quality responses based on established evaluation metrics, the responses rarely capture participants' specific motives or personalized examples, and thus perform poorly with respect to richness. We further find low inter-rater reliability between LLMs and humans in the assessment of both quality and richness metrics. Our study offers a cautionary tale for scaling and evaluating qualitative research with LLMs.

📄 View on arXiv 🌐 View on ar5iv 📑 PDF 🎉 Report Code Found

Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

📜 Similar Papers

In the same crypt — Human-Computer Interaction

R.I.P. 👻 Ghosted

Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy

Ben Shneiderman

cs.HC 🏛 International journal of human computer interactions 📚 974 cites 6 years ago

R.I.P. 👻 Ghosted

Improving fairness in machine learning systems: What do industry practitioners need?

Kenneth Holstein, Jennifer Wortman Vaughan, ... (+3 more)

cs.HC 🏛 CHI 📚 919 cites 7 years ago

R.I.P. 👻 Ghosted

Identifying Stable Patterns over Time for Emotion Recognition from EEG

Wei-Long Zheng, Jia-Yi Zhu, Bao-Liang Lu

cs.HC 🏛 IEEE TAC 📚 837 cites 10 years ago

R.I.P. 👻 Ghosted

Questioning the AI: Informing Design Practices for Explainable AI User Experiences

Q. Vera Liao, Daniel Gruen, Sarah Miller

cs.HC 🏛 CHI 📚 835 cites 6 years ago

R.I.P. 👻 Ghosted

Deep Learning for Sensor-based Human Activity Recognition: Overview, Challenges and Opportunities

Kaixuan Chen, Dalin Zhang, ... (+4 more)

cs.HC 🏛 ACM CSUR 📚 788 cites 6 years ago

R.I.P. 👻 Ghosted

Educational data mining and learning analytics: An updated survey

C. Romero, S. Ventura

cs.HC 🏛 WIREs Data Mining Knowl. Discov. 📚 787 cites 2 years ago

Died the same way — 👻 Ghosted

R.I.P. 👻 Ghosted

Federated Learning: Strategies for Improving Communication Efficiency

Jakub Konečný, H. Brendan McMahan, ... (+4 more)

cs.LG 🏛 arXiv 📚 5.2K cites 9 years ago

R.I.P. 👻 Ghosted

In-Datacenter Performance Analysis of a Tensor Processing Unit

Norman P. Jouppi, Cliff Young, ... (+73 more)

cs.AR 🏛 ISCA 📚 5.1K cites 9 years ago

R.I.P. 👻 Ghosted

Deep Convolutional Neural Networks for Computer-Aided Detection: CNN Architectures, Dataset Characteristics and Transfer Learning

Hoo-Chang Shin, Holger R. Roth, ... (+7 more)

cs.CV 🏛 IEEE TMI 📚 4.9K cites 10 years ago

R.I.P. 👻 Ghosted

Explanation in Artificial Intelligence: Insights from the Social Sciences

Tim Miller

cs.AI 🏛 AI 📚 4.9K cites 9 years ago