CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model

August 13, 2026 ยท Grace Period ยท ๐Ÿ› ICASSP 2027

โณ Grace Period
This paper is less than 90 days old. We give authors time to release their code before passing judgment.
Authors Nhan Phan, Ilona Lรคhteenmรคki, Anna von Zansen, Olli-Pekka Pauna, Yaroslav Getman, Tamรกs Grรณsz, Mikko Kurimo arXiv ID 2608.13101 Category cs.CL: Computation & Language Cross-listed eess.AS Citations 0 Venue ICASSP 2027
Abstract
Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large language models to assess learners' speaking performance. However, existing studies provide limited analysis of how acoustic and content information contribute to predictions and how stable the resulting performance is. We propose CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content. On the Speak & Improve Corpus 2025, CASA achieves a root mean square error (RMSE) of 0.358, improving on the previous best RMSE while using approximately half the estimated inference parameters. The general-purpose architecture is designed for adaptation to other ASA corpora without structural changes and relies on three handcrafted fluency features. Through ablations and repeated runs, we analyze the individual and complementary contributions of acoustic and content information, examine performance variability, and demonstrate the potential of large language model reasoning for training-free content validation.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

๐Ÿ“œ Similar Papers

In the same crypt โ€” Computation & Language

๐ŸŒ… ๐ŸŒ… Old Age

Attention Is All You Need

Ashish Vaswani, Noam Shazeer, ... (+6 more)

cs.CL ๐Ÿ› NeurIPS ๐Ÿ“š 166.0K cites 9 years ago