| 501 |
Toward Corpus Size Requirements for Training and Evaluating Depression Risk Models Using Spoken Language
Tomek Rutowski, Amir Harati, ... (+4 more)
|
👻
Ghosted
|
cs.CL
|
11 |
1 year ago |
| 502 |
Complex-Valued Time-Frequency Self-Attention for Speech Dereverberation
Vinay Kothapally, John H. L. Hansen
|
👻
Ghosted
|
eess.AS
|
11 |
3 years ago |
| 503 |
Bayesian Networks for the robust and unbiased prediction of depression and its symptoms utilizing speech and multimodal data
Salvatore Fara, Orlaith Hickey, ... (+4 more)
|
👻
Ghosted
|
cs.LG
|
11 |
3 years ago |
| 504 |
VCSE: Time-Domain Visual-Contextual Speaker Extraction Network
Junjie Li, Meng Ge, ... (+3 more)
|
👻
Ghosted
|
cs.CV
|
11 |
3 years ago |
| 505 |
Unsupervised domain adaptation for speech recognition with unsupervised error correction
Long Mai, Julie Carson-Berndsen
|
👻
Ghosted
|
cs.SD
|
11 |
3 years ago |
| 506 |
Minimum Latency Training of Sequence Transducers for Streaming End-to-End Speech Recognition
Yusuke Shinohara, Shinji Watanabe
|
👻
Ghosted
|
eess.AS
|
11 |
3 years ago |
| 507 |
BERT, can HE predict contrastive focus? Predicting and controlling prominence in neural TTS using a language model
Brooke Stephenson, Laurent Besacier, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
11 |
4 years ago |
| 508 |
QbyE-MLPMixer: Query-by-Example Open-Vocabulary Keyword Spotting using MLPMixer
Jinmiao Huang, Waseem Gharbieh, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
11 |
4 years ago |
| 509 |
Automatic Prosody Annotation with Pre-Trained Text-Speech Model
Ziqian Dai, Jianwei Yu, ... (+6 more)
|
👻
Ghosted
|
cs.SD
|
11 |
4 years ago |
| 510 |
Extracting Targeted Training Data from ASR Models, and How to Mitigate It
Ehsan Amid, Om Thakkar, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
11 |
4 years ago |
| 511 |
MUSE: Flexible Voiceprint Receptive Fields and Multi-Path Fusion Enhanced Taylor Transformer for U-Net-based Speech Enhancement
Zizhen Lin, Xiaoting Chen, Junyu Wang
|
👻
Ghosted
|
cs.SD
|
10 |
2 years ago |
| 512 |
Leveraging speaker attribute information using multi task learning for speaker verification and diarization
Chau Luu, Peter Bell, Steve Renals
|
👻
Ghosted
|
cs.SD
|
10 |
5 years ago |
| 513 |
Style Attuned Pre-training and Parameter Efficient Fine-tuning for Spoken Language Understanding
Jin Cao, Jun Wang, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
10 |
5 years ago |
| 514 |
Direct multimodal few-shot learning of speech and images
Leanne Nortje, Herman Kamper
|
👻
Ghosted
|
cs.CL
|
10 |
5 years ago |
| 515 |
Prototypical Q Networks for Automatic Conversational Diagnosis and Few-Shot New Disease Adaption
Hongyin Luo, Shang-Wen Li, James Glass
|
👻
Ghosted
|
cs.CL
|
10 |
6 years ago |
| 516 |
Large scale weakly and semi-supervised learning for low-resource video ASR
Kritika Singh, Vimal Manohar, ... (+8 more)
|
👻
Ghosted
|
eess.AS
|
10 |
6 years ago |
| 517 |
Neural Zero-Inflated Quality Estimation Model For Automatic Speech Recognition System
Kai Fan, Jiayi Wang, ... (+5 more)
|
👻
Ghosted
|
cs.CL
|
10 |
6 years ago |
| 518 |
Identifying Personality Traits Using Overlap Dynamics in Multiparty Dialogue
Mingzhi Yu, Emer Gilmartin, Diane Litman
|
👻
Ghosted
|
cs.CL
|
10 |
6 years ago |
| 519 |
Neural MultiVoice Models for Expressing Novel Personalities in Dialog
Shereen Oraby, Lena Reed, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
10 |
7 years ago |
| 520 |
Multimodal speech synthesis architecture for unsupervised speaker adaptation
Hieu-Thi Luong, Junichi Yamagishi
|
👻
Ghosted
|
eess.AS
|
10 |
7 years ago |
| 521 |
Learning weakly supervised multimodal phoneme embeddings
Rahma Chaabouni, Ewan Dunbar, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
10 |
9 years ago |
| 522 |
Articulation rate in Swedish child-directed speech increases as a function of the age of the child even when surprisal is controlled for
Johan Sjons, Thomas Hörberg, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
10 |
9 years ago |
| 523 |
NN-grams: Unifying neural network and n-gram language models for Speech Recognition
Babak Damavandi, Shankar Kumar, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
10 |
10 years ago |
| 524 |
A Nonparametric Bayesian Approach for Spoken Term detection by Example Query
Amir Hossein Harati Nejad Torbati, Joseph Picone
|
👻
Ghosted
|
cs.CL
|
10 |
10 years ago |
| 525 |
Incremental Blockwise Beam Search for Simultaneous Speech Translation with Controllable Quality-Latency Tradeoff
Peter Polák, Brian Yan, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
10 |
2 years ago |
| 526 |
A Multitask Training Approach to Enhance Whisper with Contextual Biasing and Open-Vocabulary Keyword Spotting
Yuang Li, Min Zhang, ... (+8 more)
|
👻
Ghosted
|
cs.AI
|
10 |
2 years ago |
| 527 |
LAMASSU: Streaming Language-Agnostic Multilingual Speech Recognition and Translation Using Neural Transducers
Peidong Wang, Eric Sun, ... (+6 more)
|
👻
Ghosted
|
cs.CL
|
10 |
3 years ago |
| 528 |
Data Augmentation for Low-Resource Quechua ASR Improvement
Rodolfo Zevallos, Nuria Bel, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
10 |
4 years ago |
| 529 |
Leveraging Acoustic Contextual Representation by Audio-textual Cross-modal Learning for Conversational ASR
Kun Wei, Yike Zhang, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
10 |
4 years ago |
| 530 |
What Does it Take to Generalize SER Model Across Datasets? A Comprehensive Benchmark
Adham Ibrahim, Shady Shehata, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
9 |
2 years ago |
| 531 |
Domain Adaptation Using Class Similarity for Robust Speech Recognition
Han Zhu, Jiangjiang Zhao, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
9 |
5 years ago |
| 532 |
Unsupervised vs. transfer learning for multimodal one-shot matching of speech and images
Leanne Nortje, Herman Kamper
|
👻
Ghosted
|
cs.CL
|
9 |
5 years ago |
| 533 |
Prosody Learning Mechanism for Speech Synthesis System Without Text Length Limit
Zhen Zeng, Jianzong Wang, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
9 |
5 years ago |
| 534 |
BlaBla: Linguistic Feature Extraction for Clinical Analysis in Multiple Languages
Abhishek Shivkumar, Jack Weston, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
9 |
6 years ago |
| 535 |
Improved low-resource Somali speech recognition by semi-supervised acoustic and language model training
Astik Biswas, Raghav Menon, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
9 |
7 years ago |
| 536 |
Multi-Graph Decoding for Code-Switching ASR
Emre Yılmaz, Samuel Cohen, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
9 |
7 years ago |
| 537 |
Tied Hidden Factors in Neural Networks for End-to-End Speaker Recognition
Antonio Miguel, Jorge Llombart, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
9 |
7 years ago |
| 538 |
Statistical Model Compression for Small-Footprint Natural Language Understanding
Grant P. Strimel, Kanthashree Mysore Sathyendra, Stanislav Peshterliev
|
👻
Ghosted
|
cs.CL
|
9 |
8 years ago |
| 539 |
Neural Language Codes for Multilingual Acoustic Models
Markus Müller, Sebastian Stüker, Alex Waibel
|
👻
Ghosted
|
cs.CL
|
9 |
8 years ago |
| 540 |
Fast and Accurate OOV Decoder on High-Level Features
Yuri Khokhlov, Natalia Tomashenko, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
9 |
9 years ago |
| 541 |
Acoustic data-driven lexicon learning based on a greedy pronunciation selection framework
Xiaohui Zhang, Vimal Manohar, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
9 |
9 years ago |
| 542 |
Interpretable Temporal Class Activation Representation for Audio Spoofing Detection
Menglu Li, Xiao-Ping Zhang
|
👻
Ghosted
|
cs.SD
|
9 |
2 years ago |
| 543 |
GSQA: An End-to-End Model for Generative Spoken Question Answering
Min-Han Shih, Ho-Lam Chung, ... (+5 more)
|
💤
Eternal Rest
|
cs.CL
|
9 |
2 years ago |
| 544 |
How to Construct Perfect and Worse-than-Coin-Flip Spoofing Countermeasures: A Word of Warning on Shortcut Learning
Hye-jin Shim, Rosa González Hautamäki, ... (+2 more)
|
👻
Ghosted
|
cs.LG
|
9 |
3 years ago |
| 545 |
Iterative autoregression: a novel trick to improve your low-latency speech enhancement model
Pavel Andreev, Nicholas Babaev, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
9 |
3 years ago |
| 546 |
Unify and Conquer: How Phonetic Feature Representation Affects Polyglot Text-To-Speech (TTS)
Ariadna Sanchez, Alessio Falai, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
9 |
4 years ago |
| 547 |
Space-Efficient Representation of Entity-centric Query Language Models
Christophe Van Gysel, Mirko Hannemann, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
9 |
4 years ago |
| 548 |
AdvEst: Adversarial Perturbation Estimation to Classify and Detect Adversarial Attacks against Speaker Identification
Sonal Joshi, Saurabh Kataria, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
9 |
4 years ago |
| 549 |
A low latency ASR-free end to end spoken language understanding system
Mohamed Mhiri, Samuel Myer, Vikrant Singh Tomar
|
👻
Ghosted
|
cs.CV
|
8 |
5 years ago |
| 550 |
Auxiliary Sequence Labeling Tasks for Disfluency Detection
Dongyub Lee, Byeongil Ko, ... (+6 more)
|
👻
Ghosted
|
cs.CL
|
8 |
5 years ago |