| 601 |
FinChat: Corpus and evaluation setup for Finnish chat conversations on everyday topics
Katri Leino, Juho Leinonen, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
6 |
5 years ago |
| 602 |
Neural PLDA Modeling for End-to-End Speaker Verification
Shreyas Ramoji, Prashant Krishnan, Sriram Ganapathy
|
🌅
Old Age
|
eess.AS
|
6 |
5 years ago |
| 603 |
Applying GPGPU to Recurrent Neural Network Language Model based Fast Network Search in the Real-Time LVCSR
Kyungmin Lee, Chiyoun Park, ... (+3 more)
|
👻
Ghosted
|
cs.CL
|
6 |
6 years ago |
| 604 |
Chirp Complex Cepstrum-based Decomposition for Asynchronous Glottal Analysis
Thomas Drugman, Thierry Dutoit
|
👻
Ghosted
|
cs.SD
|
6 |
6 years ago |
| 605 |
The phonetic bases of vocal expressed emotion: natural versus acted
Hira Dhamyal, Shahan Ali Memon, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
6 |
6 years ago |
| 606 |
Multi-lingual Dialogue Act Recognition with Deep Learning Methods
Jiří Martínek, Pavel Král, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
6 |
7 years ago |
| 607 |
Large-Scale Mixed-Bandwidth Deep Neural Network Acoustic Modeling for Automatic Speech Recognition
Khoi-Nguyen C. Mac, Xiaodong Cui, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
6 |
7 years ago |
| 608 |
NIESR: Nuisance Invariant End-to-end Speech Recognition
I-Hung Hsu, Ayush Jaiswal, Premkumar Natarajan
|
👻
Ghosted
|
cs.CL
|
6 |
7 years ago |
| 609 |
Attention model for articulatory features detection
Ievgen Karaulov, Dmytro Tkanov
|
👻
Ghosted
|
eess.AS
|
6 |
7 years ago |
| 610 |
Code-Switching Detection Using ASR-Generated Language Posteriors
Qinyi Wang, Emre Yılmaz, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
6 |
7 years ago |
| 611 |
Neural Named Entity Recognition from Subword Units
Abdalghani Abujabal, Judith Gaspers
|
👻
Ghosted
|
cs.CL
|
6 |
7 years ago |
| 612 |
State Gradients for RNN Memory Analysis
Lyan Verwimp, Hugo Van hamme, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
6 |
8 years ago |
| 613 |
Large-scale Speaker Retrieval on Random Speaker Variability Subspace
Suwon Shon, Younggun Lee, Taesu Kim
|
👻
Ghosted
|
eess.AS
|
6 |
7 years ago |
| 614 |
A Batch Noise Contrastive Estimation Approach for Training Large Vocabulary Language Models
Youssef Oualil, Dietrich Klakow
|
👻
Ghosted
|
cs.CL
|
6 |
8 years ago |
| 615 |
Sequential Recurrent Neural Networks for Language Modeling
Youssef Oualil, Clayton Greenberg, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
6 |
9 years ago |
| 616 |
Audio Content based Geotagging in Multimedia
Anurag Kumar, Benjamin Elizalde, Bhiksha Raj
|
👻
Ghosted
|
cs.SD
|
6 |
10 years ago |
| 617 |
DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
Heng-Jui Chang, Hongyu Gong, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
6 |
1 year ago |
| 618 |
FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
Zhongweiyang Xu, Ali Aroudi, ... (+5 more)
|
👻
Ghosted
|
cs.SD
|
6 |
1 year ago |
| 619 |
HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
Xinlei Niu, Jing Zhang, Charles Patrick Martin
|
👻
Ghosted
|
cs.SD
|
6 |
2 years ago |
| 620 |
Relationship between auditory and semantic entrainment using Deep Neural Networks (DNN)
Jay Kejriwal, Štefan Beňuš
|
👻
Ghosted
|
cs.CL
|
6 |
2 years ago |
| 621 |
Ontology-aware Learning and Evaluation for Audio Tagging
Haohe Liu, Qiuqiang Kong, ... (+4 more)
|
💤
Eternal Rest
|
eess.AS
|
6 |
3 years ago |
| 622 |
CCATMos: Convolutional Context-aware Transformer Network for Non-intrusive Speech Quality Assessment
Yuchen Liu, Li-Chia Yang, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
6 |
3 years ago |
| 623 |
Analysis of Self-Attention Head Diversity for Conformer-based Automatic Speech Recognition
Kartik Audhkhasi, Yinghui Huang, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
6 |
3 years ago |
| 624 |
A Study of Modeling Rising Intonation in Cantonese Neural Speech Synthesis
Qibing Bai, Tom Ko, Yu Zhang
|
👻
Ghosted
|
eess.AS
|
6 |
3 years ago |
| 625 |
Unsupervised Speaker Diarization that is Agnostic to Language, Overlap-Aware, and Tuning Free
M. Iftekhar Tanveer, Diego Casabuena, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
6 |
3 years ago |
| 626 |
PoeticTTS -- Controllable Poetry Reading for Literary Studies
Julia Koch, Florian Lux, ... (+7 more)
|
👻
Ghosted
|
eess.AS
|
6 |
4 years ago |
| 627 |
GlowVC: Mel-spectrogram space disentangling model for language-independent text-free voice conversion
Magdalena Proszewska, Grzegorz Beringer, ... (+4 more)
|
👻
Ghosted
|
eess.AS
|
6 |
4 years ago |
| 628 |
Toward Low-Cost End-to-End Spoken Language Understanding
Marco Dinarelli, Marco Naguib, François Portet
|
👻
Ghosted
|
cs.CL
|
6 |
4 years ago |
| 629 |
Nonwords Pronunciation Classification in Language Development Tests for Preschool Children
Ilja Baumann, Dominik Wagner, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
6 |
4 years ago |
| 630 |
Accelerating Inference and Language Model Fusion of Recurrent Neural Network Transducers via End-to-End 4-bit Quantization
Andrea Fasoli, Chia-Yu Chen, ... (+6 more)
|
👻
Ghosted
|
cs.CL
|
6 |
4 years ago |
| 631 |
The Emotion is Not One-hot Encoding: Learning with Grayscale Label for Emotion Recognition in Conversation
Joosung Lee
|
👻
Ghosted
|
cs.CL
|
6 |
4 years ago |
| 632 |
End-to-end speech recognition modeling from de-identified data
Martin Flechl, Shou-Chun Yin, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
6 |
4 years ago |
| 633 |
Human-in-the-loop Speaker Adaptation for DNN-based Multi-speaker TTS
Kenta Udagawa, Yuki Saito, Hiroshi Saruwatari
|
👻
Ghosted
|
cs.SD
|
6 |
4 years ago |
| 634 |
Device-Directed Speech Detection: Regularization via Distillation for Weakly-Supervised Models
Vineet Garg, Ognjen Rudovic, ... (+6 more)
|
👻
Ghosted
|
eess.AS
|
6 |
4 years ago |
| 635 |
Comparative Analysis of Personalized Voice Activity Detection Systems: Assessing Real-World Effectiveness
Satyam Kumar, Sai Srujana Buddi, ... (+7 more)
|
👻
Ghosted
|
eess.AS
|
5 |
2 years ago |
| 636 |
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
Shivam Mehta, Harm Lameris, ... (+4 more)
|
👻
Ghosted
|
eess.AS
|
5 |
2 years ago |
| 637 |
From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology
Haoyang Li, Yuchen Hu, ... (+4 more)
|
👻
Ghosted
|
eess.AS
|
5 |
1 year ago |
| 638 |
Efficient Multimodal Neural Networks for Trigger-less Voice Assistants
Sai Srujana Buddi, Utkarsh Oggy Sarawgi, ... (+5 more)
|
👻
Ghosted
|
cs.LG
|
5 |
3 years ago |
| 639 |
A Snoring Sound Dataset for Body Position Recognition: Collection, Annotation, and Analysis
Li Xiao, Xiuping Yang, ... (+7 more)
|
👻
Ghosted
|
cs.SD
|
5 |
2 years ago |
| 640 |
Knowledge Distillation for Singing Voice Detection
Soumava Paul, Gurunath Reddy M, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
5 |
5 years ago |
| 641 |
Perceptimatic: A human speech perception benchmark for unsupervised subword modelling
Juliette Millet, Ewan Dunbar
|
👻
Ghosted
|
cs.CL
|
5 |
5 years ago |
| 642 |
RECOApy: Data recording, pre-processing and phonetic transcription for end-to-end speech-based applications
Adriana Stan
|
👻
Ghosted
|
eess.AS
|
5 |
5 years ago |
| 643 |
Do face masks introduce bias in speech technologies? The case of automated scoring of speaking proficiency
Anastassia Loukina, Keelan Evanini, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
5 |
5 years ago |
| 644 |
Data balancing for boosting performance of low-frequency classes in Spoken Language Understanding
Judith Gaspers, Quynh Do, Fabian Triefenbach
|
👻
Ghosted
|
eess.AS
|
5 |
5 years ago |
| 645 |
Speaker Re-identification with Speaker Dependent Speech Enhancement
Yanpei Shi, Qiang Huang, Thomas Hain
|
👻
Ghosted
|
eess.AS
|
5 |
6 years ago |
| 646 |
Bandwidth Embeddings for Mixed-bandwidth Speech Recognition
Gautam Mantena, Ozlem Kalinli, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
5 |
6 years ago |
| 647 |
Latent Dirichlet Allocation Based Acoustic Data Selection for Automatic Speech Recognition
Mortaza, Doulaty, Thomas Hain
|
👻
Ghosted
|
cs.CL
|
5 |
7 years ago |
| 648 |
Ultrasound tongue imaging for diarization and alignment of child speech therapy sessions
Manuel Sam Ribeiro, Aciel Eshky, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
5 |
7 years ago |
| 649 |
Synchronising audio and ultrasound by learning cross-modal embeddings
Aciel Eshky, Manuel Sam Ribeiro, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
5 |
7 years ago |
| 650 |
Memory Time Span in LSTMs for Multi-Speaker Source Separation
Jeroen Zegers, Hugo Van hamme
|
👻
Ghosted
|
cs.LG
|
5 |
7 years ago |