| 451 |
End-to-end Speech-to-Punctuated-Text Recognition
Jumon Nozaki, Tatsuya Kawahara, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
14 |
4 years ago |
| 452 |
End-to-End Text-to-Speech Based on Latent Representation of Speaking Styles Using Spontaneous Dialogue
Kentaro Mitsui, Tianyu Zhao, ... (+4 more)
|
👻
Ghosted
|
eess.AS
|
14 |
4 years ago |
| 453 |
Decoupled Federated Learning for ASR with Non-IID Data
Han Zhu, Jindong Wang, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
14 |
4 years ago |
| 454 |
Analyzing the Quality and Stability of a Streaming End-to-End On-Device Speech Recognizer
Yuan Shangguan, Kate Knister, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
13 |
6 years ago |
| 455 |
Towards Universal Dialogue Act Tagging for Task-Oriented Dialogues
Shachi Paul, Rahul Goel, Dilek Hakkani-Tür
|
👻
Ghosted
|
cs.CL
|
13 |
7 years ago |
| 456 |
Robust Spoken Language Understanding via Paraphrasing
Avik Ray, Yilin Shen, Hongxia Jin
|
👻
Ghosted
|
cs.CL
|
13 |
7 years ago |
| 457 |
Single-channel Speech Dereverberation via Generative Adversarial Training
Chenxing Li, Tieqiang Wang, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
13 |
8 years ago |
| 458 |
Sampling-based speech parameter generation using moment-matching networks
Shinnosuke Takamichi, Tomoki Koriyama, Hiroshi Saruwatari
|
👻
Ghosted
|
cs.SD
|
13 |
9 years ago |
| 459 |
Unsupervised Domain Discovery using Latent Dirichlet Allocation for Acoustic Modelling in Speech Recognition
Mortaza Doulaty, Oscar Saz, Thomas Hain
|
👻
Ghosted
|
cs.CL
|
13 |
10 years ago |
| 460 |
MMSpeech: Multi-modal Multi-task Encoder-Decoder Pre-training for Speech Recognition
Xiaohuan Zhou, Jiaming Wang, ... (+5 more)
|
👻
Ghosted
|
cs.MM
|
13 |
3 years ago |
| 461 |
4D ASR: Joint modeling of CTC, Attention, Transducer, and Mask-Predict decoders
Yui Sudo, Muhammad Shakeel, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
13 |
3 years ago |
| 462 |
A Language Agnostic Multilingual Streaming On-Device ASR System
Bo Li, Tara N. Sainath, ... (+10 more)
|
👻
Ghosted
|
eess.AS
|
13 |
3 years ago |
| 463 |
ASR Error Correction with Constrained Decoding on Operation Prediction
Jingyuan Yang, Rongjun Li, Wei Peng
|
👻
Ghosted
|
cs.CL
|
13 |
3 years ago |
| 464 |
Improving Transformer-based Conversational ASR by Inter-Sentential Attention Mechanism
Kun Wei, Pengcheng Guo, Ning Jiang
|
👻
Ghosted
|
cs.SD
|
13 |
4 years ago |
| 465 |
Transfer Learning for Robust Low-Resource Children's Speech ASR with Transformers and Source-Filter Warping
Jenthe Thienpondt, Kris Demuynck
|
👻
Ghosted
|
eess.AS
|
13 |
4 years ago |
| 466 |
Parameter-Efficient Learning for Text-to-Speech Accent Adaptation
Li-Jen Yang, Chao-Han Huck Yang, Jen-Tzung Chien
|
👻
Ghosted
|
cs.SD
|
12 |
3 years ago |
| 467 |
Bootstrap an end-to-end ASR system by multilingual training, transfer learning, text-to-text mapping and synthetic audio
Manuel Giollo, Deniz Gunceler, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
12 |
5 years ago |
| 468 |
Sequence-to-Sequence Learning via Attention Transfer for Incremental Speech Recognition
Sashi Novitasari, Andros Tjandra, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
12 |
5 years ago |
| 469 |
Learning Explicit Prosody Models and Deep Speaker Embeddings for Atypical Voice Conversion
Disong Wang, Songxiang Liu, ... (+4 more)
|
👻
Ghosted
|
eess.AS
|
12 |
5 years ago |
| 470 |
On Minimum Word Error Rate Training of the Hybrid Autoregressive Transducer
Liang Lu, Zhong Meng, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
12 |
5 years ago |
| 471 |
Audio Dequantization for High Fidelity Audio Generation in Flow-based Neural Vocoder
Hyun-Wook Yoon, Sang-Hoon Lee, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
12 |
5 years ago |
| 472 |
Relational Teacher Student Learning with Neural Label Embedding for Device Adaptation in Acoustic Scene Classification
Hu Hu, Sabato Marco Siniscalchi, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
12 |
5 years ago |
| 473 |
Exploring Deep Hybrid Tensor-to-Vector Network Architectures for Regression Based Speech Enhancement
Jun Qi, Hu Hu, ... (+4 more)
|
👻
Ghosted
|
eess.AS
|
12 |
5 years ago |
| 474 |
Early Stage LM Integration Using Local and Global Log-Linear Combination
Wilfried Michel, Ralf Schlüter, Hermann Ney
|
👻
Ghosted
|
eess.AS
|
12 |
6 years ago |
| 475 |
Predicting Behavior in Cancer-Afflicted Patient and Spouse Interactions using Speech and Language
Sandeep Nallan Chakravarthula, Haoqi Li, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
12 |
6 years ago |
| 476 |
A computational model of early language acquisition from audiovisual experiences of young infants
Okko Räsänen, Khazar Khorrami
|
👻
Ghosted
|
cs.CL
|
12 |
7 years ago |
| 477 |
On the Contributions of Visual and Textual Supervision in Low-Resource Semantic Speech Retrieval
Ankita Pasad, Bowen Shi, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
12 |
7 years ago |
| 478 |
Self-imitating Feedback Generation Using GAN for Computer-Assisted Pronunciation Training
Seung Hee Yang, Minhwa Chung
|
👻
Ghosted
|
cs.CL
|
12 |
7 years ago |
| 479 |
Enriching Rare Word Representations in Neural Language Models by Embedding Matrix Augmentation
Yerbolat Khassanov, Zhiping Zeng, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
12 |
7 years ago |
| 480 |
Semi-supervised and Active-learning Scenarios: Efficient Acoustic Model Refinement for a Low Resource Indian Language
Maharajan Chellapriyadharshini, Anoop Toffy, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
12 |
7 years ago |
| 481 |
Gated Recurrent Unit Based Acoustic Modeling with Future Context
Jie Li, Xiaorui Wang, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
12 |
8 years ago |
| 482 |
Automatic Speech Recognition and Topic Identification for Almost-Zero-Resource Languages
Matthew Wiesner, Chunxi Liu, ... (+7 more)
|
👻
Ghosted
|
cs.CL
|
12 |
8 years ago |
| 483 |
The IBM Speaker Recognition System: Recent Advances and Error Analysis
Seyed Omid Sadjadi, Jason Pelecanos, Sriram Ganapathy
|
👻
Ghosted
|
cs.CL
|
12 |
10 years ago |
| 484 |
Leveraging Word Embeddings for Spoken Document Summarization
Kuan-Yu Chen, Shih-Hung Liu, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
12 |
11 years ago |
| 485 |
Zero-Shot Fake Video Detection by Audio-Visual Consistency
Xiaolou Li, Zehua Liu, ... (+4 more)
|
👻
Ghosted
|
cs.SD
|
12 |
2 years ago |
| 486 |
Contextual Biasing with the Knuth-Morris-Pratt Matching Algorithm
Weiran Wang, Zelin Wu, ... (+11 more)
|
👻
Ghosted
|
cs.CL
|
12 |
2 years ago |
| 487 |
Visual Transformers for Primates Classification and Covid Detection
Steffen Illium, Robert Müller, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
12 |
3 years ago |
| 488 |
When Is TTS Augmentation Through a Pivot Language Useful?
Nathaniel Robinson, Perez Ogayo, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
12 |
4 years ago |
| 489 |
Improving Deliberation by Text-Only and Semi-Supervised Training
Ke Hu, Tara N. Sainath, ... (+5 more)
|
👻
Ghosted
|
cs.CL
|
12 |
4 years ago |
| 490 |
Towards End-to-End Private Automatic Speaker Recognition
Francisco Teixeira, Alberto Abad, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
12 |
4 years ago |
| 491 |
Detecting Unintended Memorization in Language-Model-Fused ASR
W. Ronny Huang, Steve Chien, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
12 |
4 years ago |
| 492 |
Few-shot Class-incremental Audio Classification Using Adaptively-refined Prototypes
Wei Xie, Yanxiong Li, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
11 |
3 years ago |
| 493 |
Compact Speaker Embedding: lrx-vector
Munir Georges, Jonathan Huang, Tobias Bocklet
|
👻
Ghosted
|
eess.AS
|
11 |
5 years ago |
| 494 |
Black-box Adaptation of ASR for Accented Speech
Kartik Khandelwal, Preethi Jyothi, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
11 |
6 years ago |
| 495 |
A non-causal FFTNet architecture for speech enhancement
Muhammed PV Shifas, Nagaraj Adiga, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
11 |
6 years ago |
| 496 |
The NTNU System at the Interspeech 2020 Non-Native Children's Speech ASR Challenge
Tien-Hong Lo, Fu-An Chao, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
11 |
6 years ago |
| 497 |
Evolutionary Algorithm Enhanced Neural Architecture Search for Text-Independent Speaker Verification
Xiaoyang Qu, Jianzong Wang, Jing Xiao
|
👻
Ghosted
|
eess.AS
|
11 |
5 years ago |
| 498 |
Keyword Spotting for Hearing Assistive Devices Robust to External Speakers
Iván López-Espejo, Zheng-Hua Tan, Jesper Jensen
|
👻
Ghosted
|
cs.SD
|
11 |
7 years ago |
| 499 |
An Empirical Analysis of the Correlation of Syntax and Prosody
Arne Köhn, Timo Baumann, Oskar Dörfler
|
👻
Ghosted
|
cs.CL
|
11 |
8 years ago |
| 500 |
Learning Speech Rate in Speech Recognition
Xiangyu Zeng, Shi Yin, Dong Wang
|
👻
Ghosted
|
cs.CL
|
11 |
11 years ago |