| 701 |
A Training and Inference Strategy Using Noisy and Enhanced Speech as Target for Speech Enhancement without Clean Speech
Li-Wei Chen, Yao-Fei Cheng, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
4 |
3 years ago |
| 702 |
Towards continually learning new languages
Ngoc-Quan Pham, Jan Niehues, Alexander Waibel
|
👻
Ghosted
|
cs.CL
|
4 |
3 years ago |
| 703 |
Deep LSTM Spoken Term Detection using Wav2Vec 2.0 Recognizer
Jan Švec, Jan Lehečka, Luboš Šmídl
|
👻
Ghosted
|
cs.CL
|
4 |
3 years ago |
| 704 |
Predicting pairwise preferences between TTS audio stimuli using parallel ratings data and anti-symmetric twin neural networks
Cassia Valentini-Botinhao, Manuel Sam Ribeiro, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
4 |
3 years ago |
| 705 |
Parameter-Efficient Conformers via Sharing Sparsely-Gated Experts for End-to-End Speech Recognition
Ye Bai, Jie Li, ... (+6 more)
|
👻
Ghosted
|
eess.AS
|
4 |
3 years ago |
| 706 |
Towards Cross-speaker Reading Style Transfer on Audiobook Dataset
Xiang Li, Changhe Song, ... (+4 more)
|
👻
Ghosted
|
cs.SD
|
4 |
3 years ago |
| 707 |
Comparison and Analysis of New Curriculum Criteria for End-to-End ASR
Georgios Karakasidis, Tamás Grósz, Mikko Kurimo
|
👻
Ghosted
|
eess.AS
|
4 |
3 years ago |
| 708 |
Benchmarking Transformers-based models on French Spoken Language Understanding tasks
Oralie Cattan, Sahar Ghannay, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
4 |
4 years ago |
| 709 |
Investigating the Impact of Cross-lingual Acoustic-Phonetic Similarities on Multilingual Speech Recognition
Muhammad Umar Farooq, Thomas Hain
|
👻
Ghosted
|
cs.CL
|
4 |
4 years ago |
| 710 |
Resource-Efficient Speech Quality Prediction through Quantization Aware Training and Binary Activation Maps
Mattias Nilsson, Riccardo Miccini, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
3 |
2 years ago |
| 711 |
ADI-20: Arabic Dialect Identification dataset and models
Haroun Elleuch, Salima Mdhaffar, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
3 |
8 months ago |
| 712 |
Zero-Shot Mono-to-Binaural Speech Synthesis
Alon Levkovitch, Julian Salazar, ... (+5 more)
|
👻
Ghosted
|
cs.SD
|
3 |
1 year ago |
| 713 |
How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
Ailin Liu, Pepijn Vunderink, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
3 |
2 years ago |
| 714 |
Voice Passing : a Non-Binary Voice Gender Prediction System for evaluating Transgender voice transition
David Doukhan, Simon Devauchelle, ... (+5 more)
|
👻
Ghosted
|
eess.AS
|
3 |
2 years ago |
| 715 |
Video Multimodal Emotion Recognition System for Real World Applications
Sun-Kyung Lee, Jong-Hwan Kim
|
👻
Ghosted
|
cs.HC
|
3 |
2 years ago |
| 716 |
Quantifying the perceptual value of lexical and non-lexical channels in speech
Sarenne Wallbridge, Peter Bell, Catherine Lai
|
👻
Ghosted
|
cs.CL
|
3 |
3 years ago |
| 717 |
AutoSpeech 2020: The Second Automated Machine Learning Challenge for Speech Classification
Jingsong Wang, Tom Ko, ... (+5 more)
|
👻
Ghosted
|
cs.AI
|
3 |
5 years ago |
| 718 |
Text Augmentation for Language Models in High Error Recognition Scenario
Karel Beneš, Lukáš Burget
|
👻
Ghosted
|
cs.CL
|
3 |
5 years ago |
| 719 |
Augmenting Images for ASR and TTS through Single-loop and Dual-loop Multimodal Chain Framework
Johanes Effendi, Andros Tjandra, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
3 |
5 years ago |
| 720 |
Complementary Language Model and Parallel Bi-LRNN for False Trigger Mitigation
Rishika Agarwal, Xiaochuan Niu, ... (+4 more)
|
👻
Ghosted
|
eess.AS
|
3 |
5 years ago |
| 721 |
"This is Houston. Say again, please". The Behavox system for the Apollo-11 Fearless Steps Challenge (phase II)
Arseniy Gorin, Daniil Kulko, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
3 |
5 years ago |
| 722 |
Evaluating Automatically Generated Phoneme Captions for Images
Justin van der Hout, Zoltán D'Haese, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
3 |
5 years ago |
| 723 |
Automatic Quality Assessment for Audio-Visual Verification Systems. The LOVe submission to NIST SRE Challenge 2019
Grigory Antipov, Nicolas Gengembre, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
3 |
5 years ago |
| 724 |
An End-to-End Audio Classification System based on Raw Waveforms and Mix-Training Strategy
Jiaxu Chen, Jing Hao, ... (+4 more)
|
👻
Ghosted
|
eess.AS
|
3 |
6 years ago |
| 725 |
Active Annotation: bootstrapping annotation lexicon and guidelines for supervised NLU learning
Federico Marinelli, Alessandra Cervone, ... (+4 more)
|
👻
Ghosted
|
cs.CL
|
3 |
6 years ago |
| 726 |
Singing voice phoneme segmentation by hierarchically inferring syllable and phoneme onset positions
Rong Gong, Xavier Serra
|
👻
Ghosted
|
cs.SD
|
3 |
8 years ago |
| 727 |
Machine Assisted Analysis of Vowel Length Contrasts in Wolof
Elodie Gauthier, Laurent Besacier, Sylvie Voisin
|
👻
Ghosted
|
cs.CL
|
3 |
9 years ago |
| 728 |
Learning Similarity Functions for Pronunciation Variations
Einat Naaman, Yossi Adi, Joseph Keshet
|
👻
Ghosted
|
cs.CL
|
3 |
9 years ago |
| 729 |
Joint Sound Source Separation and Speaker Recognition
Jeroen Zegers, Hugo Van hamme
|
👻
Ghosted
|
cs.SD
|
3 |
10 years ago |
| 730 |
A real-time framework for visual feedback of articulatory data using statistical shape models
Kristy James, Alexander Hewer, ... (+2 more)
|
👻
Ghosted
|
cs.HC
|
3 |
9 years ago |
| 731 |
Plagiarism Detection in Polyphonic Music using Monaural Signal Separation
Soham De, Indradyumna Roy, ... (+5 more)
|
👻
Ghosted
|
cs.SD
|
3 |
11 years ago |
| 732 |
Modular Speech-to-Text Translation for Zero-Shot Cross-Modal Transfer
Paul-Ambroise Duquenne, Holger Schwenk, Benoît Sagot
|
👻
Ghosted
|
cs.CL
|
3 |
2 years ago |
| 733 |
Generating Multilingual Gender-Ambiguous Text-to-Speech Voices
Konstantinos Markopoulos, Georgia Maniati, ... (+10 more)
|
👻
Ghosted
|
cs.SD
|
3 |
3 years ago |
| 734 |
Extending Compositional Attention Networks for Social Reasoning in Videos
Christina Sartzetaki, Georgios Paraskevopoulos, Alexandros Potamianos
|
👻
Ghosted
|
cs.CV
|
3 |
3 years ago |
| 735 |
Blank Collapse: Compressing CTC emission for the faster decoding
Minkyu Jung, Ohhyeok Kwon, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
3 |
3 years ago |
| 736 |
Random Utterance Concatenation Based Data Augmentation for Improving Short-video Speech Recognition
Yist Y. Lin, Tao Han, ... (+7 more)
|
👻
Ghosted
|
eess.AS
|
3 |
3 years ago |
| 737 |
Joint Speech Translation and Named Entity Recognition
Marco Gaido, Sara Papi, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
3 |
3 years ago |
| 738 |
Pronunciation Modeling of Foreign Words for Mandarin ASR by Considering the Effect of Language Transfer
Lei Wang, Rong Tong
|
👻
Ghosted
|
cs.CL
|
3 |
3 years ago |
| 739 |
From Disfluency Detection to Intent Detection and Slot Filling
Mai Hoang Dao, Thinh Hung Truong, Dat Quoc Nguyen
|
👻
Ghosted
|
cs.CL
|
3 |
3 years ago |
| 740 |
Integrating Form and Meaning: A Multi-Task Learning Model for Acoustic Word Embeddings
Badr M. Abdullah, Bernd Möbius, Dietrich Klakow
|
👻
Ghosted
|
cs.CL
|
3 |
3 years ago |
| 741 |
Enhancing Semantic Understanding with Self-supervised Methods for Abstractive Dialogue Summarization
Hyunjae Lee, Jaewoong Yun, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
3 |
3 years ago |
| 742 |
VQ-T: RNN Transducers using Vector-Quantized Prediction Network States
Jiatong Shi, George Saon, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
3 |
3 years ago |
| 743 |
Multiple-hypothesis RNN-T Loss for Unsupervised Fine-tuning and Self-training of Neural Transducer
Cong-Thanh Do, Mohan Li, Rama Doddipatla
|
👻
Ghosted
|
cs.CL
|
3 |
3 years ago |
| 744 |
Non-Linear Pairwise Language Mappings for Low-Resource Multilingual Acoustic Model Fusion
Muhammad Umar Farooq, Darshan Adiga Haniya Narayana, Thomas Hain
|
👻
Ghosted
|
cs.CL
|
3 |
4 years ago |
| 745 |
A Temporal Extension of Latent Dirichlet Allocation for Unsupervised Acoustic Unit Discovery
Werner van der Merwe, Herman Kamper, Johan du Preez
|
👻
Ghosted
|
eess.AS
|
3 |
4 years ago |
| 746 |
Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval
Ruofan Hu, Yan Xia, ... (+6 more)
|
👻
Ghosted
|
cs.IR
|
2 |
1 year ago |
| 747 |
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
Youjun Chen, Xurong Xie, ... (+7 more)
|
👻
Ghosted
|
cs.SD
|
2 |
1 year ago |
| 748 |
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction
Wenxuan Wu, Shuai Wang, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
2 |
1 year ago |
| 749 |
PIAVE: A Pose-Invariant Audio-Visual Speaker Extraction Network
Qinghua Liu, Meng Ge, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
2 |
2 years ago |
| 750 |
Masked Proxy Loss For Text-Independent Speaker Verification
Jiachen Lian, Aiswarya Vinod Kumar, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
2 |
5 years ago |