| 551 |
Analysis of Disfluency in Children's Speech
Trang Tran, Morgan Tinkler, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
8 |
5 years ago |
| 552 |
Generative Adversarial Training Data Adaptation for Very Low-resource Automatic Speech Recognition
Kohei Matsuura, Masato Mimura, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
8 |
6 years ago |
| 553 |
Semi-supervised Learning for Multi-speaker Text-to-speech Synthesis Using Discrete Speech Representation
Tao Tu, Yuan-Jui Chen, ... (+2 more)
|
🌅
Old Age
|
eess.AS
|
8 |
6 years ago |
| 554 |
Exploring TTS without T Using Biologically/Psychologically Motivated Neural Network Modules (ZeroSpeech 2020)
Takashi Morita, Hiroki Koda
|
👻
Ghosted
|
cs.CL
|
8 |
6 years ago |
| 555 |
CNN-LSTM models for Multi-Speaker Source Separation using Bayesian Hyper Parameter Optimization
Jeroen Zegers, Hugo Van hamme
|
👻
Ghosted
|
cs.LG
|
8 |
6 years ago |
| 556 |
Self-supervised pre-training with acoustic configurations for replay spoofing detection
Hye-jin Shim, Hee-Soo Heo, ... (+2 more)
|
👻
Ghosted
|
cs.LG
|
8 |
6 years ago |
| 557 |
Modeling user context for valence prediction from narratives
Aniruddha Tammewar, Alessandra Cervone, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
8 |
7 years ago |
| 558 |
Iterative Delexicalization for Improved Spoken Language Understanding
Avik Ray, Yilin Shen, Hongxia Jin
|
👻
Ghosted
|
cs.CL
|
8 |
6 years ago |
| 559 |
Empirical Evaluation of Sequence-to-Sequence Models for Word Discovery in Low-resource Settings
Marcely Zanon Boito, Aline Villavicencio, Laurent Besacier
|
👻
Ghosted
|
cs.CL
|
8 |
7 years ago |
| 560 |
Lattice-based lightly-supervised acoustic model training
Joachim Fainberg, Ondřej Klejch, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
8 |
7 years ago |
| 561 |
Audio Classification of Bit-Representation Waveform
Masaki Okawa, Takuya Saito, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
8 |
7 years ago |
| 562 |
Kernel Machines Beat Deep Neural Networks on Mask-based Single-channel Speech Enhancement
Like Hui, Siyuan Ma, Mikhail Belkin
|
👻
Ghosted
|
cs.LG
|
8 |
7 years ago |
| 563 |
Global SNR Estimation of Speech Signals using Entropy and Uncertainty Estimates from Dropout Networks
Rohith Aralikatti, Dilip Margam, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
8 |
8 years ago |
| 564 |
Waveform to Single Sinusoid Regression to Estimate the F0 Contour from Noisy Speech Using Recurrent Deep Neural Networks
Akihiro Kato, Tomi Kinnunen
|
👻
Ghosted
|
eess.AS
|
8 |
8 years ago |
| 565 |
Opinion Dynamics Modeling for Movie Review Transcripts Classification with Hidden Conditional Random Fields
Valentin Barriere, Chloé Clavel, Slim Essid
|
👻
Ghosted
|
cs.CL
|
8 |
8 years ago |
| 566 |
GMM-Free Flat Start Sequence-Discriminative DNN Training
Gábor Gosztolya, Tamás Grósz, László Tóth
|
👻
Ghosted
|
cs.CL
|
8 |
9 years ago |
| 567 |
Prompt Tuning for Audio Deepfake Detection: Computationally Efficient Test-time Domain Adaptation with Limited Target Dataset
Hideyuki Oiso, Yuto Matsunaga, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
8 |
1 year ago |
| 568 |
Exploring In-Context Learning of Textless Speech Language Model for Speech Classification Tasks
Ming-Hao Hsu, Kai-Wei Chang, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
8 |
2 years ago |
| 569 |
SlothSpeech: Denial-of-service Attack Against Speech Recognition Models
Mirazul Haque, Rutvij Shah, ... (+4 more)
|
👻
Ghosted
|
cs.SD
|
8 |
3 years ago |
| 570 |
Explicit Intensity Control for Accented Text-to-speech
Rui Liu, Haolin Zuo, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
8 |
3 years ago |
| 571 |
Compute Cost Amortized Transformer for Streaming ASR
Yi Xie, Jonathan Macoskey, ... (+6 more)
|
👻
Ghosted
|
cs.CL
|
8 |
4 years ago |
| 572 |
Speech Emotion: Investigating Model Representations, Multi-Task Learning and Knowledge Distillation
Vikramjit Mitra, Hsiang-Yun Sherry Chien, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
8 |
4 years ago |
| 573 |
A Multi-Task BERT Model for Schema-Guided Dialogue State Tracking
Eleftherios Kapelonis, Efthymios Georgiou, Alexandros Potamianos
|
👻
Ghosted
|
cs.CL
|
8 |
4 years ago |
| 574 |
Bottleneck Low-rank Transformers for Low-resource Spoken Language Understanding
Pu Wang, Hugo Van hamme
|
👻
Ghosted
|
cs.CL
|
8 |
4 years ago |
| 575 |
Acoustic Modeling for End-to-End Empathetic Dialogue Speech Synthesis Using Linguistic and Prosodic Contexts of Dialogue History
Yuto Nishimura, Yuki Saito, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
8 |
4 years ago |
| 576 |
DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval
Yifei Xin, Xuxin Cheng, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
7 |
1 year ago |
| 577 |
Enhancing Speech-Driven 3D Facial Animation with Audio-Visual Guidance from Lip Reading Expert
Han EunGi, Oh Hyun-Bin, ... (+5 more)
|
👻
Ghosted
|
cs.CV
|
7 |
2 years ago |
| 578 |
Utterance-Wise Meeting Transcription System Using Asynchronous Distributed Microphones
Shota Horiguchi, Yusuke Fujita, Kenji Nagamatsu
|
👻
Ghosted
|
eess.AS
|
7 |
5 years ago |
| 579 |
CTC-synchronous Training for Monotonic Attention Model
Hirofumi Inaguma, Masato Mimura, Tatsuya Kawahara
|
👻
Ghosted
|
cs.CL
|
7 |
6 years ago |
| 580 |
Practical applicability of deep neural networks for overlapping speaker separation
Pieter Appeltans, Jeroen Zegers, Hugo Van hamme
|
👻
Ghosted
|
cs.LG
|
7 |
6 years ago |
| 581 |
SANTLR: Speech Annotation Toolkit for Low Resource Languages
Xinjian Li, Zhong Zhou, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
7 |
6 years ago |
| 582 |
Comparison of Lattice-Free and Lattice-Based Sequence Discriminative Training Criteria for LVCSR
Wilfried Michel, Ralf Schlüter, Hermann Ney
|
👻
Ghosted
|
eess.AS
|
7 |
7 years ago |
| 583 |
Modeling Interpersonal Influence of Verbal Behavior in Couples Therapy Dyadic Interactions
Sandeep Nallan Chakravarthula, Brian Baucom, Panayiotis Georgiou
|
👻
Ghosted
|
cs.CL
|
7 |
8 years ago |
| 584 |
A Generative Model for Score Normalization in Speaker Recognition
Albert Swart, Niko Brummer
|
👻
Ghosted
|
stat.ML
|
7 |
8 years ago |
| 585 |
Contrastive Entropy: A new evaluation metric for unnormalized language models
Kushal Arora, Anand Rangarajan
|
👻
Ghosted
|
cs.CL
|
7 |
10 years ago |
| 586 |
RT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement
Honglie Chen, Rodrigo Mira, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
7 |
2 years ago |
| 587 |
Missingness-resilient Video-enhanced Multimodal Disfluency Detection
Payal Mohapatra, Shamika Likhite, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
7 |
2 years ago |
| 588 |
Single-channel speech enhancement using learnable loss mixup
Oscar Chang, Dung N. Tran, Kazuhito Koishida
|
👻
Ghosted
|
eess.AS
|
7 |
2 years ago |
| 589 |
Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion
Yun Chen, Lingxiao Yang, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
7 |
2 years ago |
| 590 |
A Compact End-to-End Model with Local and Global Context for Spoken Language Identification
Fei Jia, Nithin Rao Koluguri, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
7 |
3 years ago |
| 591 |
Spoken Term Detection and Relevance Score Estimation using Dot-Product of Pronunciation Embeddings
Jan Švec, Luboš Šmídl, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
7 |
3 years ago |
| 592 |
ESSumm: Extractive Speech Summarization from Untranscribed Meeting
Jun Wang
|
👻
Ghosted
|
eess.AS
|
7 |
3 years ago |
| 593 |
Thutmose Tagger: Single-pass neural model for Inverse Text Normalization
Alexandra Antonova, Evelina Bakhturina, Boris Ginsburg
|
👻
Ghosted
|
cs.CL
|
7 |
3 years ago |
| 594 |
Improving Data Driven Inverse Text Normalization using Data Augmentation
Laxmi Pandey, Debjyoti Paul, ... (+7 more)
|
👻
Ghosted
|
cs.CL
|
7 |
4 years ago |
| 595 |
ASR-Generated Text for Language Model Pre-training Applied to Speech Tasks
Valentin Pelloin, Franck Dary, ... (+5 more)
|
👻
Ghosted
|
cs.CL
|
7 |
4 years ago |
| 596 |
Global-Local Convolution with Spiking Neural Networks for Energy-efficient Keyword Spotting
Shuai Wang, Dehao Zhang, ... (+5 more)
|
👻
Ghosted
|
cs.SD
|
6 |
2 years ago |
| 597 |
Text Injection for Neural Contextual Biasing
Zhong Meng, Zelin Wu, ... (+6 more)
|
👻
Ghosted
|
cs.CL
|
6 |
2 years ago |
| 598 |
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach
Umberto Cappellazzo, Minsu Kim, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
6 |
1 year ago |
| 599 |
Reduce and Reconstruct: ASR for Low-Resource Phonetic Languages
Anuj Diwan, Preethi Jyothi
|
👻
Ghosted
|
eess.AS
|
6 |
5 years ago |
| 600 |
TMT: A Transformer-based Modal Translator for Improving Multimodal Sequence Representations in Audio Visual Scene-aware Dialog
Wubo Li, Dongwei Jiang, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
6 |
5 years ago |