| 251 |
COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning
Jing Pan, Jian Wu, ... (+5 more)
|
👻
Ghosted
|
cs.CL
|
33 |
2 years ago |
| 252 |
Are disentangled representations all you need to build speaker anonymization systems?
Pierre Champion, Denis Jouvet, Anthony Larcher
|
👻
Ghosted
|
cs.SD
|
33 |
3 years ago |
| 253 |
Evaluating the reliability of acoustic speech embeddings
Robin Algayres, Mohamed Salah Zaiem, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
32 |
5 years ago |
| 254 |
Understanding Self-Attention of Self-Supervised Audio Transformers
Shu-wen Yang, Andy T. Liu, Hung-yi Lee
|
👻
Ghosted
|
cs.CL
|
32 |
6 years ago |
| 255 |
Dynamic Prosody Generation for Speech Synthesis using Linguistics-Driven Acoustic Embedding Selection
Shubhi Tyagi, Marco Nicolis, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
32 |
6 years ago |
| 256 |
Predicting the Leading Political Ideology of YouTube Channels Using Acoustic, Textual, and Metadata Information
Yoan Dinkov, Ahmed Ali, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
32 |
6 years ago |
| 257 |
Prosodic Phrase Alignment for Machine Dubbing
Alp Öktem, Mireia Farrús, Antonio Bonafonte
|
👻
Ghosted
|
cs.CL
|
32 |
6 years ago |
| 258 |
Spatial Pyramid Encoding with Convex Length Normalization for Text-Independent Speaker Verification
Youngmoon Jung, Younggwan Kim, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
32 |
7 years ago |
| 259 |
Learning to adapt: a meta-learning approach for speaker adaptation
Ondřej Klejch, Joachim Fainberg, Peter Bell
|
👻
Ghosted
|
cs.CL
|
32 |
7 years ago |
| 260 |
Building a Unified Code-Switching ASR System for South African Languages
Emre Yılmaz, Astik Biswas, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
32 |
7 years ago |
| 261 |
Learning Acoustic Word Embeddings with Temporal Context for Query-by-Example Speech Search
Yougen Yuan, Cheung-Chi Leung, ... (+4 more)
|
👻
Ghosted
|
cs.CL
|
32 |
8 years ago |
| 262 |
Contaminated speech training methods for robust DNN-HMM distant speech recognition
Mirco Ravanelli, Maurizio Omologo
|
👻
Ghosted
|
eess.AS
|
32 |
8 years ago |
| 263 |
ASR2K: Speech Recognition for Around 2000 Languages without Audio
Xinjian Li, Florian Metze, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
32 |
3 years ago |
| 264 |
Pretrained Semantic Speech Embeddings for End-to-End Spoken Language Understanding via Cross-Modal Teacher-Student Learning
Pavel Denisov, Ngoc Thang Vu
|
👻
Ghosted
|
eess.AS
|
31 |
6 years ago |
| 265 |
Exploiting Syntactic Features in a Parsed Tree to Improve End-to-End TTS
Haohan Guo, Frank K. Soong, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
31 |
7 years ago |
| 266 |
Adversarial Audio: A New Information Hiding Method and Backdoor for DNN-based Speech Recognition Models
Yehao Kong, Jiliang Zhang
|
👻
Ghosted
|
cs.CR
|
31 |
7 years ago |
| 267 |
Hide and Speak: Towards Deep Neural Networks for Speech Steganography
Felix Kreuk, Yossi Adi, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
31 |
7 years ago |
| 268 |
SPEECH-COCO: 600k Visually Grounded Spoken Captions Aligned to MSCOCO Data Set
William Havard, Laurent Besacier, Olivier Rosec
|
👻
Ghosted
|
cs.CL
|
31 |
8 years ago |
| 269 |
Unsupervised End-to-End Learning of Discrete Linguistic Units for Voice Conversion
Andy T. Liu, Po-chun Hsu, Hung-yi Lee
|
👻
Ghosted
|
cs.CL
|
30 |
7 years ago |
| 270 |
Building a mixed-lingual neural TTS system with only monolingual data
Liumeng Xue, Wei Song, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
30 |
7 years ago |
| 271 |
Spoken Language Intent Detection using Confusion2Vec
Prashanth Gurunath Shivakumar, Mu Yang, Panayiotis Georgiou
|
👻
Ghosted
|
cs.CL
|
30 |
7 years ago |
| 272 |
Fast ASR-free and almost zero-resource keyword spotting using DTW and CNNs for humanitarian monitoring
Raghav Menon, Herman Kamper, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
30 |
8 years ago |
| 273 |
Anti-Spoofing Using Transfer Learning with Variational Information Bottleneck
Youngsik Eom, Yeonghyeon Lee, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
30 |
4 years ago |
| 274 |
Multitask Training with Text Data for End-to-End Speech Recognition
Peidong Wang, Tara N. Sainath, Ron J. Weiss
|
👻
Ghosted
|
cs.CL
|
29 |
5 years ago |
| 275 |
A Convolutional Deep Markov Model for Unsupervised Speech Representation Learning
Sameer Khurana, Antoine Laurent, ... (+5 more)
|
👻
Ghosted
|
eess.AS
|
29 |
6 years ago |
| 276 |
Conversational Emotion Analysis via Attention Mechanisms
Zheng Lian, Jianhua Tao, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
29 |
6 years ago |
| 277 |
Unsupervised Word Segmentation from Speech with Attention
Pierre Godard, Marcely Zanon-Boito, ... (+5 more)
|
👻
Ghosted
|
cs.CL
|
29 |
8 years ago |
| 278 |
Leveraging translations for speech transcription in low-resource settings
Antonis Anastasopoulos, David Chiang
|
👻
Ghosted
|
cs.CL
|
29 |
8 years ago |
| 279 |
Diff-E: Diffusion-based Learning for Decoding Imagined Speech EEG
Soowon Kim, Young-Eun Lee, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
28 |
2 years ago |
| 280 |
End-to-End Spoken Language Understanding Without Full Transcripts
Hong-Kwang J. Kuo, Zoltán Tüske, ... (+8 more)
|
👻
Ghosted
|
cs.CL
|
28 |
5 years ago |
| 281 |
Affective Conditioning on Hierarchical Networks applied to Depression Detection from Transcribed Clinical Interviews
D. Xezonaki, G. Paraskevopoulos, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
28 |
6 years ago |
| 282 |
End-to-End Multi-Speaker Speech Recognition using Speaker Embeddings and Transfer Learning
Pavel Denisov, Ngoc Thang Vu
|
👻
Ghosted
|
eess.AS
|
28 |
6 years ago |
| 283 |
Polyphone Disambiguation for Mandarin Chinese Using Conditional Neural Network with Multi-level Embedding Features
Zexin Cai, Yaogen Yang, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
28 |
7 years ago |
| 284 |
Sampling strategies in Siamese Networks for unsupervised speech representation learning
Rachid Riad, Corentin Dancette, ... (+4 more)
|
👻
Ghosted
|
cs.CL
|
28 |
8 years ago |
| 285 |
Unspeech: Unsupervised Speech Context Embeddings
Benjamin Milde, Chris Biemann
|
👻
Ghosted
|
cs.SD
|
28 |
8 years ago |
| 286 |
Attentive Sequence-to-Sequence Learning for Diacritic Restoration of Yorùbá Language Text
Iroro Orife
|
👻
Ghosted
|
cs.CL
|
28 |
8 years ago |
| 287 |
Unsupervised Discovery of Recurring Speech Patterns Using Probabilistic Adaptive Metrics
Okko Räsänen, María Andrea Cruz Blandón
|
👻
Ghosted
|
eess.AS
|
27 |
5 years ago |
| 288 |
Cross-Lingual Speaker Verification with Domain-Balanced Hard Prototype Mining and Language-Dependent Score Normalization
Jenthe Thienpondt, Brecht Desplanques, Kris Demuynck
|
👻
Ghosted
|
eess.AS
|
27 |
6 years ago |
| 289 |
Speech to Text Adaptation: Towards an Efficient Cross-Modal Distillation
Won Ik Cho, Donghyun Kwak, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
27 |
6 years ago |
| 290 |
Who Needs Words? Lexicon-Free Speech Recognition
Tatiana Likhomanenko, Gabriel Synnaeve, Ronan Collobert
|
👻
Ghosted
|
cs.CL
|
27 |
7 years ago |
| 291 |
An Ensemble of Transfer, Semi-supervised and Supervised Learning Methods for Pathological Heart Sound Classification
Ahmed Imtiaz Humayun, Md. Tauhiduzzaman Khan, ... (+3 more)
|
👻
Ghosted
|
cs.CV
|
27 |
8 years ago |
| 292 |
UTD-CRSS Systems for 2016 NIST Speaker Recognition Evaluation
Chunlei Zhang, Fahimeh Bahmaninezhad, ... (+4 more)
|
👻
Ghosted
|
cs.CL
|
27 |
9 years ago |
| 293 |
Visually-Aware Audio Captioning With Adaptive Audio-Visual Attention
Xubo Liu, Qiushi Huang, ... (+11 more)
|
👻
Ghosted
|
eess.AS
|
27 |
3 years ago |
| 294 |
Multi-Accent Adaptation based on Gate Mechanism
Han Zhu, Li Wang, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
26 |
5 years ago |
| 295 |
Interpretable Deep Learning Model for the Detection and Reconstruction of Dysarthric Speech
Daniel Korzekwa, Roberto Barra-Chicote, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
26 |
7 years ago |
| 296 |
Listen, Attend, Spell and Adapt: Speaker Adapted Sequence-to-Sequence ASR
Felix Weninger, Jesús Andrés-Ferrer, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
26 |
7 years ago |
| 297 |
Scalable Multi Corpora Neural Language Models for ASR
Anirudh Raju, Denis Filimonov, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
26 |
7 years ago |
| 298 |
Acoustic Modeling for Automatic Lyrics-to-Audio Alignment
Chitralekha Gupta, Emre Yılmaz, Haizhou Li
|
👻
Ghosted
|
eess.AS
|
26 |
7 years ago |
| 299 |
Objective Assessment of Social Skills Using Automated Language Analysis for Identification of Schizophrenia and Bipolar Disorder
Rohit Voleti, Stephanie Woolridge, ... (+4 more)
|
👻
Ghosted
|
cs.CL
|
26 |
7 years ago |
| 300 |
STC Speaker Recognition Systems for the VOiCES From a Distance Challenge
Sergey Novoselov, Aleksei Gusev, ... (+6 more)
|
👻
Ghosted
|
cs.SD
|
26 |
7 years ago |