| 51 |
Fast Speech Foundation Model Distillation Using Interleaved Stacking
Eungbeom Kim, Kyogu Lee
|
|
eess.AS
|
0 |
1 month ago |
| 52 |
UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction
Sangmin Lee, Eekgyun Ahn, ... (+2 more)
|
|
cs.CL
|
0 |
1 month ago |
| 53 |
SpAArSIST: Sparsified AASIST for Efficient and Reliable Anti-Spoofing
Anton Firc, Vojtěch Staněk, ... (+3 more)
|
|
cs.SD
|
0 |
1 month ago |
| 54 |
Pretrained self-supervised speech models can recognize unseen consonants
Chihiro Taguchi, Éric Le Ferrand, ... (+5 more)
|
|
cs.CL
|
0 |
1 month ago |
| 55 |
Gumbel-BEARD: Automatic Layer Selection for Self-Supervised Adaptation of Whisper in Low-Resource Domains
Zilai Wang, Natarajan Balaji Shankar, ... (+3 more)
|
|
eess.AS
|
0 |
1 month ago |
| 56 |
What Do Deepfake Speech Detectors Actually Hear?
Vojtěch Staněk, Veronika Jirmusová, ... (+4 more)
|
|
cs.SD
|
0 |
1 month ago |
| 57 |
Ethical and Technical Limits of Deepfake Speech Datasets
Vojtěch Staněk, Eva Trnovská, ... (+2 more)
|
|
cs.SD
|
0 |
1 month ago |
| 58 |
RAT: Reference-Augmented Training for ASV Anti-Spoofing
Vojtěch Staněk, Anton Firc, ... (+2 more)
|
|
cs.SD
|
0 |
1 month ago |
| 59 |
Massive Open-Vocabulary Keyword Spotting
Leonor Barreiros, Raul Monteiro, ... (+2 more)
|
|
eess.AS
|
0 |
1 month ago |
| 60 |
Multilingual Word-Level Forced Alignment with Self-Supervised Representations and Learned Dynamic Programming
Roy Weber, Meidan Zehavi, ... (+2 more)
|
|
cs.CL
|
0 |
1 month ago |
| 61 |
ANCHOR: Autoregressive Non-intrusive Chunk-Ordered Refinement for Joint Multi-Resolution Speech Quality Modeling
Zhuoyan Tao, Jiatong Shi, ... (+2 more)
|
|
eess.AS
|
0 |
1 month ago |
| 62 |
On the Effect of Segmentation Width and Cluster Size on Speech Resynthesis and Continuation in Generative Spoken Language Models
Shunsuke Kando, Wataru Nakata, ... (+2 more)
|
|
cs.CL
|
0 |
29 days ago |
| 63 |
Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS
Seymanur Akti, Alexander Waibel
|
|
cs.SD
|
0 |
29 days ago |
| 64 |
From Text Metrics to Model Internals: A Study of Whisper ASR Hallucination Detection
Jan Jasiński, Mateusz Barański, ... (+3 more)
|
|
cs.SD
|
0 |
29 days ago |
| 65 |
HALAS: A Human-Annotated Dataset of Hallucinations of Modern ASR Systems
Mateusz Barański, Jan Jasiński, ... (+3 more)
|
|
cs.SD
|
0 |
29 days ago |
| 66 |
Cross-lingual Retrieval-Augmented Classification for Dysarthria Severity Assessment
Taeyoung Jeong, Insung Lee, ... (+2 more)
|
|
cs.SD
|
0 |
29 days ago |
| 67 |
Learning to Evade: Adaptive Attacks on Audio Watermarking
Weikang Ding, Hanqing Guo, ... (+5 more)
|
|
cs.SD
|
0 |
1 month ago |
| 68 |
What Do Neural Networks Learn for TDOA Estimation? A Cross-Architecture Probing Study
Yaozhong Kang, Jiang Wang, ... (+4 more)
|
|
cs.SD
|
0 |
1 month ago |
| 69 |
Benchmarking Large Language Models for Grapheme-to-Phoneme Conversion: A Japanese Case Study
Tomoki Koriyama
|
|
cs.CL
|
0 |
1 month ago |
| 70 |
Adding Robust Code-Switching Capabilities to High Performance Multilingual ASR
Enes Yavuz Ugan, Alexander Waibel
|
|
cs.CL
|
0 |
1 month ago |
| 71 |
Integrating Facial Generation into Full-Duplex Spoken Dialogue Systems
Jingjing Jiang, Atsumoto Ohashi, Ryuichiro Higashinaka
|
|
cs.HC
|
0 |
1 month ago |
| 72 |
Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead
Muyang Du, Jason Roche, Junjie Lai
|
|
cs.SD
|
0 |
1 month ago |
| 73 |
Time-Frequency Weighted Losses for Phoneme Reconstruction in DNN-Based Speech Enhancement
Nasser-Eddine Monir, Paul Magron, Romain Serizel
|
|
cs.SD
|
0 |
1 month ago |
| 74 |
Post-Training Speech Enhancement Language Models with Perceptual Rewards
Frédéric Berdoz, Luca A. Lanzendörfer, ... (+2 more)
|
|
cs.LG
|
0 |
1 month ago |
| 75 |
Sexualised synthetic personas encode and amplify gendered power asymmetries through voice
Alice Ross, Ariadna Sanchez, ... (+3 more)
|
|
eess.AS
|
0 |
1 month ago |
| 76 |
An Evaluation Framework for Text-to-Speech Voice Reconstruction
Ariadna Sanchez, Christoph Minixhofer, ... (+4 more)
|
|
eess.AS
|
0 |
1 month ago |
| 77 |
Synthetic Audio Generation Framework for Air Traffic Control Speech Recognition
Raphaël Bagat, Zhe Zhang, ... (+3 more)
|
|
cs.CL
|
0 |
1 month ago |
| 78 |
Towards Dys-XAI: Influence-Based Explanations for Dysarthria Severity Assessment
Xiaoliang Wu, Qiyang Sun, ... (+4 more)
|
|
cs.AI
|
0 |
1 month ago |
| 79 |
LISE : Listenable Interpretable Speaker Embeddings
Xiaoliang Wu, Chongxin Gan, ... (+3 more)
|
|
cs.SD
|
0 |
1 month ago |
| 80 |
Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach
Tzu-Chieh Wei, Yi-Cheng Lin, ... (+5 more)
|
|
eess.AS
|
0 |
1 month ago |
| 81 |
LLM-Based Multi-Reference Evaluation for Efficient and Robust Assessment of Phrase Break Annotations
Younghan Park, Hoyeon Lee, ... (+2 more)
|
|
cs.CL
|
0 |
1 month ago |
| 82 |
Imitation Learning for Elder-Facing Speech Synthesis
Dongrui Han, Weidong Chen, ... (+4 more)
|
|
cs.SD
|
0 |
1 month ago |
| 83 |
Exploiting Neural Audio Codec Latents for Adversarial Audio Attacks
Sameek Bhattacharya, Bharath Krishnamurthy, Ajita Rattani
|
|
cs.SD
|
0 |
1 month ago |
| 84 |
Repurposing a Speech Classifier for Guided Diffusion-Based Speech Generation
Rostislav Makarov, Timo Gerkmann
|
|
eess.AS
|
0 |
1 month ago |
| 85 |
PASQA: Pitch-Accent-Focused Speech Quality Assessment Model Trained on Synthetic Speech with Accent Errors
Masaya Kawamura, Yuma Shirahata, ... (+2 more)
|
|
eess.AS
|
0 |
1 month ago |
| 86 |
Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations
Masato Takagi, Masaya Kawamura, ... (+2 more)
|
|
eess.AS
|
0 |
1 month ago |
| 87 |
Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal
Syeda Faiza Ahmed Sara, Shammur Absar Chowdhury
|
|
cs.CL
|
0 |
1 month ago |
| 88 |
Low-Burden Data Augmentation for Dysarthric ASR via Zero-Shot Voice Cloning
Satwinder Singh, Qianli Wang, ... (+5 more)
|
|
eess.AS
|
0 |
1 month ago |
| 89 |
PrefSQA: Pairwise Preference Prediction for Speech Quality Assessment and the Critical Role of High Quality Datasets
Junyi Fan, Donald S. Williamson
|
|
cs.SD
|
0 |
1 month ago |
| 90 |
IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages
Sakshi Joshi, Dhruv Subhash Rathi, ... (+5 more)
|
|
eess.AS
|
0 |
1 month ago |
| 91 |
Adaptive Speech-to-Spike Encoding for Spiking Neural Networks
Taharim Rahman Anon, Jakaria Islam Emon
|
|
cs.NE
|
0 |
1 month ago |
| 92 |
Mitigating Scoring Errors and Compensating for Nonverbal Subtests in Speech-Based Dementia Assessment
Franziska Braun, Christopher Witzl, ... (+5 more)
|
|
eess.AS
|
0 |
1 month ago |
| 93 |
QC-GAN: A Parameter-Efficient Quaternion Conformer GAN for High-Fidelity Speech Enhancement
Shogo Yamauchi, Hideaki Tamori, ... (+3 more)
|
|
cs.SD
|
0 |
1 month ago |
| 94 |
Fair Cognitive Impairment Detection Through Unlearning
William Nguyen, Jiali Cheng, Hadi Amiri
|
|
cs.LG
|
0 |
1 month ago |
| 95 |
MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data
Subhankar Ghosh, Jason Li, ... (+5 more)
|
|
cs.SD
|
0 |
1 month ago |
| 96 |
Learning task-specific subspaces via interventional post-training of speech foundation models
Jack Cox, Jon Barker
|
|
cs.CL
|
0 |
1 month ago |
| 97 |
A Generalized Formalism of Auto-Regressive Decoding for Speech Processing
Julia Gachot, Philipp Allgeuer, ... (+2 more)
|
|
cs.SD
|
0 |
1 month ago |
| 98 |
Perceptual compensation for tonal context in self-supervised speech models
James Kirby, Ioana Krehan, Michele Gubian
|
|
cs.CL
|
0 |
1 month ago |
| 99 |
When Multiple Scripts Matter: Evaluating ASR in Clinical Settings
Jean Seo, Minkyu Kim, ... (+4 more)
|
|
cs.CL
|
0 |
1 month ago |
| 100 |
Towards Speech Impairment Prediction in German-Speaking Individuals with Amyotrophic Lateral Sclerosis
Monica Gonzalez-Machorro, Ricarda von Heynitz, ... (+6 more)
|
|
cs.HC
|
0 |
1 month ago |