| 1 |
TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis
Xi Wang, Jie Wang, ... (+9 more)
|
|
cs.CL
|
0 |
2 months ago |
| 2 |
Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models
Sandra Arcos-Holzinger, Sarah M. Erfani, ... (+2 more)
|
|
eess.AS
|
0 |
2 months ago |
| 3 |
Multilingual Phonological Feature Recognition with Self-Supervised Speech Models
Abner Hernandez, Tomás Arias-Vergara, ... (+3 more)
|
|
cs.CL
|
0 |
1 month ago |
| 4 |
MeCo: One-Step MeanFlow-based Corrector for Multi-Channel Speech Separation
Dohwan Kim, Jung-Woo Choi
|
|
eess.AS
|
0 |
1 month ago |
| 5 |
Overcoming Decoder Inconsistencies in Whisper for Dravidian and Low-Resource Languages
Chowdam Venkata Kumar, Kumud Tripathi, Pankaj Wasnik
|
|
cs.CL
|
0 |
1 month ago |
| 6 |
A Finetuned SpeechLLM for Joint Multi-Granular L2 Assessment and Natural-Language Rationales
Aditya Kamlesh Parikh, Cristian Tejedor-Garcia, ... (+2 more)
|
|
cs.CL
|
0 |
1 month ago |
| 7 |
Few-shot Class-variable Incremental Audio Classification via Prototype Adaptation and Pseudo Class-variable Training
Yanxiong Li, Guoqing Chen, ... (+2 more)
|
|
eess.AS
|
0 |
1 month ago |
| 8 |
Segment-level Tree Search for Long Meeting Document Summarization
Sangwon Ryu, Heejin Do, ... (+5 more)
|
|
cs.CL
|
0 |
1 month ago |
| 9 |
TinyGiantALM: A Compact Audio-Language Model for Intent-Aware Reasoning under Resource Constraints
Vinh-Thuan Ly
|
|
cs.SD
|
0 |
1 month ago |
| 10 |
Paediatric-HGNN: A Hybrid Heterogeneous Graph Neural Network for Detecting Disfluency in Children's Speech via Multiscale Acoustic Fusion
Rashini Liyanarachchi, Rachael Mackay, ... (+3 more)
|
|
eess.AS
|
0 |
1 month ago |
| 11 |
The Lipreading Gap: Do VSR Models Perceive Visual Speech Like Human Lipreaders?
Rishabh Jain, Naomi Harte
|
|
cs.CV
|
0 |
1 month ago |
| 12 |
Contrastive Training with LLM-generated Near-Misses for Robust Code-Switching Speech Recognition
Tung X. Nguyen, Hieu Minh Truong, ... (+4 more)
|
|
cs.CL
|
0 |
1 month ago |
| 13 |
SEAM: Shortcut-Aware Real-Time Detection of Scripted vs. Spontaneous Speech for Interview Guardrails
Vsevolod, Kovalev, Pranay Manocha
|
|
eess.AS
|
0 |
1 month ago |
| 14 |
HybridCodec: Fast Dual-Stream, Semantically Enhanced Neural Audio Codec
Arjun Gangwar, S Umesh
|
|
cs.SD
|
0 |
1 month ago |
| 15 |
Multilingual Multi-Speaker Unit Vocoders: A Systematic Analysis of Discrete Speech Representations
Naman Kothari, Arjun Gangwar, ... (+2 more)
|
|
cs.SD
|
0 |
1 month ago |
| 16 |
USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding
Heng-Jui Chang, Alexander H. Liu, ... (+5 more)
|
|
eess.AS
|
0 |
1 month ago |
| 17 |
A Hierarchical Feature Engineering Framework for Automated Classification of Phonotraumatic and Non-Phonotraumatic Vocal Hyperfunction
June-Woo Kim, Kangwook Jang, ... (+2 more)
|
|
cs.SD
|
0 |
1 month ago |
| 18 |
ProSarc: Prosody-Aware Sarcasm Recognition Framework via Temporal Prosodic Incongruity
Prathamjyot Singh, Ashima Sood, ... (+2 more)
|
|
cs.AI
|
0 |
1 month ago |
| 19 |
To Be Multimodal or Not to Be: Query-Adaptive Audio-Visual Person Retrieval via Active Modality Detection
Erfan Loweimi, Mengjie Qian, ... (+8 more)
|
|
cs.CL
|
0 |
1 month ago |
| 20 |
Domain-Aware Mispronunciation Detection and Diagnosis Using Language-Specific Statistical Graphs
Huu Tuong Tu, Hanh Nguyen, ... (+4 more)
|
|
cs.CL
|
0 |
1 month ago |
| 21 |
Read What You Hear: Reference-Free Hypotheses Evaluation with Acoustic Discrepancy
Zhihan Li, Hankun Wang, ... (+4 more)
|
|
eess.AS
|
0 |
1 month ago |
| 22 |
Probing Low Frame Rate Degradation in Neural Audio Codecs
Alex Gichamba, Moise Busogi
|
|
cs.SD
|
0 |
1 month ago |
| 23 |
ArtNet: A JEPA-Like Articulatory Predictive Framework for Robust Zero-Shot Phoneme Recognition
Zeqian Hu, Fuliang Weng, ... (+2 more)
|
|
cs.SD
|
0 |
1 month ago |
| 24 |
Dual-Granularity Orthogonal Disentanglement for Generalizable Audio Deepfake Detection
Zhuodong Liu, Hugen Lv, ... (+2 more)
|
|
cs.SD
|
0 |
1 month ago |
| 25 |
From Awareness to Adherence: Bridging the Context Gap in Spoken Dialogue Systems via Context-Aware Decoding
Che Hyun Lee, Heeseung Kim, Sungroh Yoon
|
|
cs.CL
|
0 |
1 month ago |
| 26 |
TMASC: Transmasculine Attitude and Speech Corpus
Sidney Wong
|
|
cs.CL
|
0 |
1 month ago |
| 27 |
XAI-Grounded Explanation Generation for Speech Deepfake Detection with Training-Free Multimodal Large Language Models
Yupei Li, Qiyang Sun, ... (+4 more)
|
|
cs.CL
|
0 |
1 month ago |
| 28 |
Scaling Human and G2P Supervision for Robust Phonetic Transcription
Alexander Metzger, Aruna Srivastava, Ruslan Mukhamedvaleev
|
|
cs.CL
|
0 |
1 month ago |
| 29 |
Geometrically Constrained Decentralized Independent Vector Analysis for Distributed Microphone Arrays
Changda Chen, Yichen Yang, ... (+5 more)
|
|
eess.AS
|
0 |
1 month ago |
| 30 |
Acoustic Prompting via Stage-wise Modulation for Few-Shot Learning in Audio Language Models
Hyebin Cho, Jaehyuk Jang, ... (+2 more)
|
|
cs.SD
|
0 |
1 month ago |
| 31 |
Evaluating and Preserving Lexical Stress in English-to-Chinese Speech-to-Speech Translation
Yuchen Song, Xi Chen, ... (+2 more)
|
|
cs.CL
|
0 |
1 month ago |
| 32 |
FreeSonic: Training-Free Temporal-Aware Decoupled Attention for Precise Audio Editing
Yuxuan Jiang, Mingyang Han, ... (+13 more)
|
|
cs.SD
|
0 |
1 month ago |
| 33 |
EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning
Siyuan Zhang, Jian Zong, ... (+8 more)
|
|
eess.AS
|
0 |
1 month ago |
| 34 |
Beyond task performance: Decoding bioacoustic embeddings with speech features
Ines Nolasco, Jules Cauzinille, ... (+9 more)
|
|
cs.LG
|
0 |
1 month ago |
| 35 |
Listening with Attention: Entropy-Guided Explainability for Transformer-Based Audio Models
Ravi Ranjan, Utkarsh Grover, ... (+2 more)
|
|
cs.SD
|
0 |
1 month ago |
| 36 |
MoDiCoL: A Modular Diagnostic Continual Learning Dataset for Robust Speech Recognition
Theresa Pekarek Rosin, Matthias Kerzel, Stefan Wermter
|
|
cs.CL
|
0 |
1 month ago |
| 37 |
Spectro-Temporal Interference Confounds Phase Encoding in Spatial Audio Foundation Models
Yuxuan Chen, Haoyuan Yu, Peize He
|
|
cs.SD
|
0 |
1 month ago |
| 38 |
Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR
Henri-Leon Kordt, Theresa Pekarek Rosin, ... (+2 more)
|
|
cs.CL
|
0 |
1 month ago |
| 39 |
FoleyGenEx: Unified Video-to-Audio Generation with Multi-Modal Control, Temporal Alignment, and Semantic Precision
Shiyao Wang, Xijuan Zeng, ... (+5 more)
|
|
cs.SD
|
0 |
1 month ago |
| 40 |
From Tokens to Faces: Investigating Discrete Speech Representations for 3D Facial Animation
Pedro Correa, Olivier Perrotin, ... (+3 more)
|
|
cs.CL
|
0 |
1 month ago |
| 41 |
Adaptive Turn-Taking for Real-time Multi-Party Voice Agents
Soumyajit Mitra, Prabhat Pandey, ... (+3 more)
|
|
eess.AS
|
0 |
1 month ago |
| 42 |
Leveraging Audio-LLMs to Filter Speech-to-Speech Training Data
Qixu Chen, Satoshi Nakamura
|
|
cs.CL
|
0 |
1 month ago |
| 43 |
Positional Encoding in the Context of Memristor-Based Analog Computation for Automatic Speech Recognition
Benedikt Hilmes, Nick Rossenbach, Ralf Schlüter
|
|
cs.LG
|
0 |
1 month ago |
| 44 |
Predicting Cognitive Load from Speech and Interaction Dynamics in Dyadic Conversations
Tahiya Chowdhury
|
|
cs.LG
|
0 |
1 month ago |
| 45 |
PiDA: Phonetically-Informed Data Augmentation for Robust Vietnamese Speech Translation
Giang Son Nguyen, Tung X. Nguyen, ... (+4 more)
|
|
cs.CL
|
0 |
1 month ago |
| 46 |
PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue
Wen Zhang, Xiaocui Yang, ... (+4 more)
|
|
cs.CL
|
0 |
1 month ago |
| 47 |
Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation
Zhen Ye, Xu Tan, ... (+11 more)
|
|
eess.AS
|
0 |
1 month ago |
| 48 |
Lung-SRAD: Spectral-Aware Regularized Audio DASS with Dual-Axis Patch-Mix Contrastive Learning for Respiratory Sound Classification
Hemansh Shridhar, Miika Toikkanen, June-Woo Kim
|
|
cs.SD
|
0 |
1 month ago |
| 49 |
Quality Adaptive Angular Margin Learning for Respiratory Sound Classification
Yoon Tae Kim, Heejoon Koo, ... (+2 more)
|
|
cs.SD
|
0 |
1 month ago |
| 50 |
Towards Data-free and Training-free Compression for Speech Foundation Models Using Parameter Clustering
Haoning Xu, Zhaoqing Li, ... (+5 more)
|
|
cs.SD
|
0 |
1 month ago |