| 1 |
MeCo: One-Step MeanFlow-based Corrector for Multi-Channel Speech Separation
Dohwan Kim, Jung-Woo Choi
|
|
eess.AS
|
0 |
2 months ago |
| 2 |
Overcoming Decoder Inconsistencies in Whisper for Dravidian and Low-Resource Languages
Chowdam Venkata Kumar, Kumud Tripathi, Pankaj Wasnik
|
|
cs.CL
|
0 |
2 months ago |
| 3 |
A Finetuned SpeechLLM for Joint Multi-Granular L2 Assessment and Natural-Language Rationales
Aditya Kamlesh Parikh, Cristian Tejedor-Garcia, ... (+2 more)
|
|
cs.CL
|
0 |
2 months ago |
| 4 |
Few-shot Class-variable Incremental Audio Classification via Prototype Adaptation and Pseudo Class-variable Training
Yanxiong Li, Guoqing Chen, ... (+2 more)
|
|
eess.AS
|
0 |
2 months ago |
| 5 |
Segment-level Tree Search for Long Meeting Document Summarization
Sangwon Ryu, Heejin Do, ... (+5 more)
|
|
cs.CL
|
0 |
2 months ago |
| 6 |
TinyGiantALM: A Compact Audio-Language Model for Intent-Aware Reasoning under Resource Constraints
Vinh-Thuan Ly
|
|
cs.SD
|
0 |
2 months ago |
| 7 |
Probing Low Frame Rate Degradation in Neural Audio Codecs
Alex Gichamba, Moise Busogi
|
|
cs.SD
|
0 |
2 months ago |
| 8 |
ArtNet: A JEPA-Like Articulatory Predictive Framework for Robust Zero-Shot Phoneme Recognition
Zeqian Hu, Fuliang Weng, ... (+2 more)
|
|
cs.SD
|
0 |
2 months ago |
| 9 |
Dual-Granularity Orthogonal Disentanglement for Generalizable Audio Deepfake Detection
Zhuodong Liu, Hugen Lv, ... (+2 more)
|
|
cs.SD
|
0 |
2 months ago |
| 10 |
From Awareness to Adherence: Bridging the Context Gap in Spoken Dialogue Systems via Context-Aware Decoding
Che Hyun Lee, Heeseung Kim, Sungroh Yoon
|
|
cs.CL
|
0 |
2 months ago |
| 11 |
TMASC: Transmasculine Attitude and Speech Corpus
Sidney Wong
|
|
cs.CL
|
0 |
2 months ago |
| 12 |
XAI-Grounded Explanation Generation for Speech Deepfake Detection with Training-Free Multimodal Large Language Models
Yupei Li, Qiyang Sun, ... (+4 more)
|
|
cs.CL
|
0 |
2 months ago |
| 13 |
Scaling Human and G2P Supervision for Robust Phonetic Transcription
Alexander Metzger, Aruna Srivastava, Ruslan Mukhamedvaleev
|
|
cs.CL
|
0 |
2 months ago |
| 14 |
Geometrically Constrained Decentralized Independent Vector Analysis for Distributed Microphone Arrays
Changda Chen, Yichen Yang, ... (+5 more)
|
|
eess.AS
|
0 |
2 months ago |
| 15 |
Acoustic Prompting via Stage-wise Modulation for Few-Shot Learning in Audio Language Models
Hyebin Cho, Jaehyuk Jang, ... (+2 more)
|
|
cs.SD
|
0 |
2 months ago |
| 16 |
Evaluating and Preserving Lexical Stress in English-to-Chinese Speech-to-Speech Translation
Yuchen Song, Xi Chen, ... (+2 more)
|
|
cs.CL
|
0 |
2 months ago |
| 17 |
FreeSonic: Training-Free Temporal-Aware Decoupled Attention for Precise Audio Editing
Yuxuan Jiang, Mingyang Han, ... (+13 more)
|
|
cs.SD
|
0 |
2 months ago |
| 18 |
EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning
Siyuan Zhang, Jian Zong, ... (+8 more)
|
|
eess.AS
|
0 |
2 months ago |
| 19 |
Beyond task performance: Decoding bioacoustic embeddings with speech features
Ines Nolasco, Jules Cauzinille, ... (+9 more)
|
|
cs.LG
|
0 |
2 months ago |
| 20 |
Listening with Attention: Entropy-Guided Explainability for Transformer-Based Audio Models
Ravi Ranjan, Utkarsh Grover, ... (+2 more)
|
|
cs.SD
|
0 |
2 months ago |
| 21 |
MoDiCoL: A Modular Diagnostic Continual Learning Dataset for Robust Speech Recognition
Theresa Pekarek Rosin, Matthias Kerzel, Stefan Wermter
|
|
cs.CL
|
0 |
2 months ago |
| 22 |
Spectro-Temporal Interference Confounds Phase Encoding in Spatial Audio Foundation Models
Yuxuan Chen, Haoyuan Yu, Peize He
|
|
cs.SD
|
0 |
2 months ago |
| 23 |
Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR
Henri-Leon Kordt, Theresa Pekarek Rosin, ... (+2 more)
|
|
cs.CL
|
0 |
2 months ago |
| 24 |
FoleyGenEx: Unified Video-to-Audio Generation with Multi-Modal Control, Temporal Alignment, and Semantic Precision
Shiyao Wang, Xijuan Zeng, ... (+5 more)
|
|
cs.SD
|
0 |
2 months ago |
| 25 |
From Tokens to Faces: Investigating Discrete Speech Representations for 3D Facial Animation
Pedro Correa, Olivier Perrotin, ... (+3 more)
|
|
cs.CL
|
0 |
2 months ago |
| 26 |
Adaptive Turn-Taking for Real-time Multi-Party Voice Agents
Soumyajit Mitra, Prabhat Pandey, ... (+3 more)
|
|
eess.AS
|
0 |
2 months ago |
| 27 |
Leveraging Audio-LLMs to Filter Speech-to-Speech Training Data
Qixu Chen, Satoshi Nakamura
|
|
cs.CL
|
0 |
2 months ago |
| 28 |
Positional Encoding in the Context of Memristor-Based Analog Computation for Automatic Speech Recognition
Benedikt Hilmes, Nick Rossenbach, Ralf Schlüter
|
|
cs.LG
|
0 |
2 months ago |
| 29 |
Predicting Cognitive Load from Speech and Interaction Dynamics in Dyadic Conversations
Tahiya Chowdhury
|
|
cs.LG
|
0 |
2 months ago |
| 30 |
PiDA: Phonetically-Informed Data Augmentation for Robust Vietnamese Speech Translation
Giang Son Nguyen, Tung X. Nguyen, ... (+4 more)
|
|
cs.CL
|
0 |
2 months ago |
| 31 |
PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue
Wen Zhang, Xiaocui Yang, ... (+4 more)
|
|
cs.CL
|
0 |
2 months ago |
| 32 |
Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation
Zhen Ye, Xu Tan, ... (+11 more)
|
|
eess.AS
|
0 |
2 months ago |
| 33 |
Lung-SRAD: Spectral-Aware Regularized Audio DASS with Dual-Axis Patch-Mix Contrastive Learning for Respiratory Sound Classification
Hemansh Shridhar, Miika Toikkanen, June-Woo Kim
|
|
cs.SD
|
0 |
2 months ago |
| 34 |
Quality Adaptive Angular Margin Learning for Respiratory Sound Classification
Yoon Tae Kim, Heejoon Koo, ... (+2 more)
|
|
cs.SD
|
0 |
2 months ago |
| 35 |
Towards Data-free and Training-free Compression for Speech Foundation Models Using Parameter Clustering
Haoning Xu, Zhaoqing Li, ... (+5 more)
|
|
cs.SD
|
0 |
2 months ago |
| 36 |
Fast Speech Foundation Model Distillation Using Interleaved Stacking
Eungbeom Kim, Kyogu Lee
|
|
eess.AS
|
0 |
2 months ago |
| 37 |
UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction
Sangmin Lee, Eekgyun Ahn, ... (+2 more)
|
|
cs.CL
|
0 |
2 months ago |
| 38 |
SpAArSIST: Sparsified AASIST for Efficient and Reliable Anti-Spoofing
Anton Firc, Vojtěch Staněk, ... (+3 more)
|
|
cs.SD
|
0 |
2 months ago |
| 39 |
Pretrained self-supervised speech models can recognize unseen consonants
Chihiro Taguchi, Éric Le Ferrand, ... (+5 more)
|
|
cs.CL
|
0 |
2 months ago |
| 40 |
Gumbel-BEARD: Automatic Layer Selection for Self-Supervised Adaptation of Whisper in Low-Resource Domains
Zilai Wang, Natarajan Balaji Shankar, ... (+3 more)
|
|
eess.AS
|
0 |
2 months ago |
| 41 |
What Do Deepfake Speech Detectors Actually Hear?
Vojtěch Staněk, Veronika Jirmusová, ... (+4 more)
|
|
cs.SD
|
0 |
2 months ago |
| 42 |
Ethical and Technical Limits of Deepfake Speech Datasets
Vojtěch Staněk, Eva Trnovská, ... (+2 more)
|
|
cs.SD
|
0 |
2 months ago |
| 43 |
RAT: Reference-Augmented Training for ASV Anti-Spoofing
Vojtěch Staněk, Anton Firc, ... (+2 more)
|
|
cs.SD
|
0 |
2 months ago |
| 44 |
Massive Open-Vocabulary Keyword Spotting
Leonor Barreiros, Raul Monteiro, ... (+2 more)
|
|
eess.AS
|
0 |
2 months ago |
| 45 |
Multilingual Word-Level Forced Alignment with Self-Supervised Representations and Learned Dynamic Programming
Roy Weber, Meidan Zehavi, ... (+2 more)
|
|
cs.CL
|
0 |
2 months ago |
| 46 |
ANCHOR: Autoregressive Non-intrusive Chunk-Ordered Refinement for Joint Multi-Resolution Speech Quality Modeling
Zhuoyan Tao, Jiatong Shi, ... (+2 more)
|
|
eess.AS
|
0 |
2 months ago |
| 47 |
On the Effect of Segmentation Width and Cluster Size on Speech Resynthesis and Continuation in Generative Spoken Language Models
Shunsuke Kando, Wataru Nakata, ... (+2 more)
|
|
cs.CL
|
0 |
2 months ago |
| 48 |
Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS
Seymanur Akti, Alexander Waibel
|
|
cs.SD
|
0 |
2 months ago |
| 49 |
From Text Metrics to Model Internals: A Study of Whisper ASR Hallucination Detection
Jan Jasiński, Mateusz Barański, ... (+3 more)
|
|
cs.SD
|
0 |
2 months ago |
| 50 |
HALAS: A Human-Annotated Dataset of Hallucinations of Modern ASR Systems
Mateusz Barański, Jan Jasiński, ... (+3 more)
|
|
cs.SD
|
0 |
2 months ago |