On LLM Wizards: Identifying Large Language Models' Behaviors for Wizard of Oz Experiments

July 10, 2024 · Declared Dead · 🏛 International Conference on Intelligent Virtual Agents

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Jingchao Fang, Nikos Arechiga, Keiichi Namaoshi, Nayeli Bravo, Candice Hogan, David A. Shamma arXiv ID 2407.08067 Category cs.HC: Human-Computer Interaction Cross-listed cs.AI Citations 6 Venue International Conference on Intelligent Virtual Agents Last Checked 4 months ago

Abstract

The Wizard of Oz (WoZ) method is a widely adopted research approach where a human Wizard ``role-plays'' a not readily available technology and interacts with participants to elicit user behaviors and probe the design space. With the growing ability for modern large language models (LLMs) to role-play, one can apply LLMs as Wizards in WoZ experiments with better scalability and lower cost than the traditional approach. However, methodological guidance on responsibly applying LLMs in WoZ experiments and a systematic evaluation of LLMs' role-playing ability are lacking. Through two LLM-powered WoZ studies, we take the first step towards identifying an experiment lifecycle for researchers to safely integrate LLMs into WoZ experiments and interpret data generated from settings that involve Wizards role-played by LLMs. We also contribute a heuristic-based evaluation framework that allows the estimation of LLMs' role-playing ability in WoZ experiments and reveals LLMs' behavior patterns at scale.

📄 View on arXiv 🌐 View on ar5iv 📑 PDF 🎉 Report Code Found

Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

📜 Similar Papers

In the same crypt — Human-Computer Interaction

R.I.P. 👻 Ghosted

Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy

Ben Shneiderman

cs.HC 🏛 International journal of human computer interactions 📚 974 cites 6 years ago

R.I.P. 👻 Ghosted

Improving fairness in machine learning systems: What do industry practitioners need?

Kenneth Holstein, Jennifer Wortman Vaughan, ... (+3 more)

cs.HC 🏛 CHI 📚 919 cites 7 years ago

R.I.P. 👻 Ghosted

Identifying Stable Patterns over Time for Emotion Recognition from EEG

Wei-Long Zheng, Jia-Yi Zhu, Bao-Liang Lu

cs.HC 🏛 IEEE TAC 📚 837 cites 10 years ago

R.I.P. 👻 Ghosted

Questioning the AI: Informing Design Practices for Explainable AI User Experiences

Q. Vera Liao, Daniel Gruen, Sarah Miller

cs.HC 🏛 CHI 📚 835 cites 6 years ago

R.I.P. 👻 Ghosted

Deep Learning for Sensor-based Human Activity Recognition: Overview, Challenges and Opportunities

Kaixuan Chen, Dalin Zhang, ... (+4 more)

cs.HC 🏛 ACM CSUR 📚 788 cites 6 years ago

R.I.P. 👻 Ghosted

Educational data mining and learning analytics: An updated survey

C. Romero, S. Ventura

cs.HC 🏛 WIREs Data Mining Knowl. Discov. 📚 787 cites 2 years ago

Died the same way — 👻 Ghosted

R.I.P. 👻 Ghosted

Federated Learning: Strategies for Improving Communication Efficiency

Jakub Konečný, H. Brendan McMahan, ... (+4 more)

cs.LG 🏛 arXiv 📚 5.2K cites 9 years ago

R.I.P. 👻 Ghosted

In-Datacenter Performance Analysis of a Tensor Processing Unit

Norman P. Jouppi, Cliff Young, ... (+73 more)

cs.AR 🏛 ISCA 📚 5.1K cites 9 years ago

R.I.P. 👻 Ghosted

Deep Convolutional Neural Networks for Computer-Aided Detection: CNN Architectures, Dataset Characteristics and Transfer Learning

Hoo-Chang Shin, Holger R. Roth, ... (+7 more)

cs.CV 🏛 IEEE TMI 📚 4.9K cites 10 years ago

R.I.P. 👻 Ghosted

Explanation in Artificial Intelligence: Insights from the Social Sciences

Tim Miller

cs.AI 🏛 AI 📚 4.9K cites 9 years ago