| 651 |
Unsupervised and Efficient Vocabulary Expansion for Recurrent Neural Network Language Models in ASR
Yerbolat Khassanov, Eng Siong Chng
|
👻
Ghosted
|
cs.CL
|
5 |
8 years ago |
| 652 |
Empirical Evaluation of Speaker Adaptation on DNN based Acoustic Model
Ke Wang, Junbo Zhang, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
5 |
8 years ago |
| 653 |
Empirical Evaluation of Parallel Training Algorithms on Acoustic Modeling
Wenpeng Li, BinBin Zhang, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
5 |
9 years ago |
| 654 |
Order-Preserving Abstractive Summarization for Spoken Content Based on Connectionist Temporal Classification
Bo-Ru Lu, Frank Shyu, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
5 |
8 years ago |
| 655 |
Recognize Foreign Low-Frequency Words with Similar Pairs
Xi Ma, Xiaoxi Wang, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
5 |
11 years ago |
| 656 |
Enhancing Modal Fusion by Alignment and Label Matching for Multimodal Emotion Recognition
Qifei Li, Yingming Gao, ... (+3 more)
|
👻
Ghosted
|
cs.MM
|
5 |
1 year ago |
| 657 |
MixRep: Hidden Representation Mixup for Low-Resource Speech Recognition
Jiamin Xie, John H. L. Hansen
|
👻
Ghosted
|
eess.AS
|
5 |
2 years ago |
| 658 |
Human Transcription Quality Improvement
Jian Gao, Hanbo Sun, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
5 |
2 years ago |
| 659 |
Multitask Learning for Low Resource Spoken Language Understanding
Quentin Meeus, Marie-Francine Moens, Hugo Van hamme
|
👻
Ghosted
|
cs.CL
|
5 |
3 years ago |
| 660 |
Technology Pipeline for Large Scale Cross-Lingual Dubbing of Lecture Videos into Multiple Indian Languages
Anusha Prakash, Arun Kumar, ... (+25 more)
|
👻
Ghosted
|
eess.AS
|
5 |
3 years ago |
| 661 |
Streaming Intended Query Detection using E2E Modeling for Continued Conversation
Shuo-yiin Chang, Guru Prakash, ... (+8 more)
|
👻
Ghosted
|
cs.CL
|
5 |
3 years ago |
| 662 |
A High-Quality and Large-Scale Dataset for English-Vietnamese Speech Translation
Linh The Nguyen, Nguyen Luong Tran, ... (+3 more)
|
📜
Death by README
|
cs.CL
|
5 |
3 years ago |
| 663 |
Act-Aware Slot-Value Predicting in Multi-Domain Dialogue State Tracking
Ruolin Su, Ting-Wei Wu, Biing-Hwang Juang
|
💤
Eternal Rest
|
cs.CL
|
5 |
3 years ago |
| 664 |
Mix and Match: An Empirical Study on Training Corpus Composition for Polyglot Text-To-Speech (TTS)
Ziyao Zhang, Alessio Falai, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
5 |
4 years ago |
| 665 |
A Polyphone BERT for Polyphone Disambiguation in Mandarin Chinese
Song Zhang, Ken Zheng, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
5 |
4 years ago |
| 666 |
Low-resource Accent Classification in Geographically-proximate Settings: A Forensic and Sociophonetics Perspective
Qingcheng Zeng, Dading Chong, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
5 |
4 years ago |
| 667 |
Self-supervised speech unit discovery from articulatory and acoustic features using VQ-VAE
Marc-Antoine Georges, Jean-Luc Schwartz, Thomas Hueber
|
👻
Ghosted
|
cs.CL
|
5 |
4 years ago |
| 668 |
AdaMS: Deep Metric Learning with Adaptive Margin and Adaptive Scale for Acoustic Word Discrimination
Myunghun Jung, Hoirin Kim
|
👻
Ghosted
|
eess.AS
|
5 |
3 years ago |
| 669 |
Towards Green ASR: Lossless 4-bit Quantization of a Hybrid TDNN System on the 300-hr Switchboard Corpus
Junhao Xu, Shoukang Hu, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
5 |
4 years ago |
| 670 |
Exploiting Context-dependent Duration Features for Voice Anonymization Attack Systems
Natalia Tomashenko, Emmanuel Vincent, Marc Tommasi
|
👻
Ghosted
|
cs.SD
|
4 |
1 year ago |
| 671 |
Rapport-Driven Virtual Agent: Rapport Building Dialogue Strategy for Improving User Experience at First Meeting
Muhammad Yeza Baihaqi, Angel García Contreras, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
4 |
2 years ago |
| 672 |
Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection
Shruti Palaskar, Oggi Rudovic, ... (+8 more)
|
👻
Ghosted
|
cs.CL
|
4 |
2 years ago |
| 673 |
STOPA: A Database of Systematic VariaTion Of DeePfake Audio for Open-Set Source Tracing and Attribution
Anton Firc, Manasi Chhibber, ... (+4 more)
|
👻
Ghosted
|
cs.SD
|
4 |
1 year ago |
| 674 |
AFL-Net: Integrating Audio, Facial, and Lip Modalities with a Two-step Cross-attention for Robust Speaker Diarization in the Wild
Yongkang Yin, Xu Li, ... (+2 more)
|
👻
Ghosted
|
cs.MM
|
4 |
2 years ago |
| 675 |
Stochastic Talking Face Generation Using Latent Distribution Matching
Ravindra Yadav, Ashish Sardana, ... (+2 more)
|
👻
Ghosted
|
cs.CV
|
4 |
5 years ago |
| 676 |
ICE-Talk: an Interface for a Controllable Expressive Talking Machine
Noé Tits, Kevin El Haddad, Thierry Dutoit
|
👻
Ghosted
|
eess.AS
|
4 |
5 years ago |
| 677 |
An Acoustic Segment Model Based Segment Unit Selection Approach to Acoustic Scene Classification with Partial Utterances
Hu Hu, Sabato Marco Siniscalchi, ... (+4 more)
|
👻
Ghosted
|
eess.AS
|
4 |
5 years ago |
| 678 |
Unsupervised Subword Modeling Using Autoregressive Pretraining and Cross-Lingual Phone-Aware Modeling
Siyuan Feng, Odette Scharenborg
|
👻
Ghosted
|
eess.AS
|
4 |
5 years ago |
| 679 |
Subword RNNLM Approximations for Out-Of-Vocabulary Keyword Search
Mittul Singh, Sami Virpioja, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
4 |
6 years ago |
| 680 |
Statistical Testing on ASR Performance via Blockwise Bootstrap
Zhe Liu, Fuchun Peng
|
👻
Ghosted
|
stat.ML
|
4 |
6 years ago |
| 681 |
Adapting a FrameNet Semantic Parser for Spoken Language Understanding Using Adversarial Learning
Gabriel Marzinotto, Geraldine Damnati, Frédéric Béchet
|
👻
Ghosted
|
cs.CL
|
4 |
6 years ago |
| 682 |
Self-Teaching Networks
Liang Lu, Eric Sun, Yifan Gong
|
👻
Ghosted
|
eess.AS
|
4 |
6 years ago |
| 683 |
Acoustic Model Optimization Based On Evolutionary Stochastic Gradient Descent with Anchors for Automatic Speech Recognition
Xiaodong Cui, Michael Picheny
|
👻
Ghosted
|
cs.CL
|
4 |
7 years ago |
| 684 |
Deep Neural Baselines for Computational Paralinguistics
Daniel Elsner, Stefan Langer, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
4 |
7 years ago |
| 685 |
End-to-end Adaptation with Backpropagation through WFST for On-device Speech Recognition System
Emiru Tsunoo, Yosuke Kashiwagi, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
4 |
7 years ago |
| 686 |
Low-Dimensional Bottleneck Features for On-Device Continuous Speech Recognition
David B. Ramsay, Kevin Kilgour, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
4 |
7 years ago |
| 687 |
Combining Natural Gradient with Hessian Free Methods for Sequence Training
Adnan Haider, P. C. Woodland
|
👻
Ghosted
|
cs.LG
|
4 |
7 years ago |
| 688 |
Semi-tied Units for Efficient Gating in LSTM and Highway Networks
Chao Zhang, Philip Woodland
|
👻
Ghosted
|
cs.CL
|
4 |
8 years ago |
| 689 |
Improved ASR for Under-Resourced Languages Through Multi-Task Learning with Acoustic Landmarks
Di He, Boon Pang Lim, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
4 |
8 years ago |
| 690 |
Efficient Segmental Cascades for Speech Recognition
Hao Tang, Weiran Wang, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
4 |
9 years ago |
| 691 |
Optimizing the role of human evaluation in LLM-based spoken document summarization systems
Margaret Kroll, Kelsey Kraus
|
👻
Ghosted
|
cs.AI
|
4 |
1 year ago |
| 692 |
Prosody-Driven Privacy-Preserving Dementia Detection
Dominika Woszczyk, Ranya Aloufi, Soteris Demetriou
|
👻
Ghosted
|
cs.SD
|
4 |
2 years ago |
| 693 |
Efficiently Train ASR Models that Memorize Less and Perform Better with Per-core Clipping
Lun Wang, Om Thakkar, ... (+4 more)
|
👻
Ghosted
|
cs.CR
|
4 |
2 years ago |
| 694 |
How Much Context Does My Attention-Based ASR System Need?
Robert Flynn, Anton Ragni
|
👻
Ghosted
|
cs.CL
|
4 |
2 years ago |
| 695 |
Wavelet Scattering Transform for Improving Generalization in Low-Resourced Spoken Language Identification
Spandan Dey, Premjeet Singh, Goutam Saha
|
👻
Ghosted
|
eess.AS
|
4 |
2 years ago |
| 696 |
Reduce, Reuse, Recycle: Is Perturbed Data better than Other Language augmentation for Low Resource Self-Supervised Speech Models
Asad Ullah, Alessandro Ragano, Andrew Hines
|
👻
Ghosted
|
eess.AS
|
4 |
2 years ago |
| 697 |
HypR: A comprehensive study for ASR hypothesis revising with a reference corpus
Yi-Wei Wang, Ke-Han Lu, Kuan-Yu Chen
|
👻
Ghosted
|
cs.CL
|
4 |
2 years ago |
| 698 |
Speaker-Aware Anti-Spoofing
Xuechen Liu, Md Sahidullah, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
4 |
3 years ago |
| 699 |
Semi-supervised learning for continuous emotional intensity controllable speech synthesis with disentangled representations
Yoori Oh, Juheon Lee, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
4 |
3 years ago |
| 700 |
Biased Self-supervised learning for ASR
Florian L. Kreyssig, Yangyang Shi, ... (+4 more)
|
👻
Ghosted
|
cs.CL
|
4 |
3 years ago |