| 651 |
Enriching Music Descriptions with a Finetuned-LLM and Metadata for Text-to-Music Retrieval
SeungHeon Doh, Minhee Lee, ... (+2 more)
|
👻
Ghosted
|
cs.SD
|
20 |
1 year ago |
| 652 |
Ms-senet: Enhancing Speech Emotion Recognition Through Multi-scale Feature Fusion With Squeeze-and-excitation Blocks
Mengbo Li, Yuanzhong Zheng, ... (+4 more)
|
👻
Ghosted
|
cs.SD
|
20 |
2 years ago |
| 653 |
Leveraging Large Language Models for Exploiting ASR Uncertainty
Pranay Dighe, Yi Su, ... (+5 more)
|
👻
Ghosted
|
cs.CL
|
20 |
2 years ago |
| 654 |
On the Tradeoff between Privacy Preservation and Byzantine-Robustness in Decentralized Learning
Haoxiang Ye, Heng Zhu, Qing Ling
|
👻
Ghosted
|
cs.LG
|
20 |
2 years ago |
| 655 |
A Joint Convolutional and Spatial Quad-Directional LSTM Network for Phase Unwrapping
Malsha V. Perera, Ashwin De Silva
|
👻
Ghosted
|
cs.LG
|
20 |
5 years ago |
| 656 |
Independent Sign Language Recognition with 3D Body, Hands, and Face Reconstruction
Agelos Kratimenos, Georgios Pavlakos, Petros Maragos
|
👻
Ghosted
|
cs.CV
|
20 |
5 years ago |
| 657 |
Adaptable Multi-Domain Language Model for Transformer ASR
Taewoo Lee, Min-Joong Lee, ... (+11 more)
|
👻
Ghosted
|
eess.AS
|
20 |
5 years ago |
| 658 |
Melody Harmonization Using Orderless NADE, Chord Balancing, and Blocked Gibbs Sampling
Chung-En Sun, Yi-Wei Chen, ... (+3 more)
|
👻
Ghosted
|
cs.SD
|
20 |
5 years ago |
| 659 |
Robust Unsupervised Audio-visual Speech Enhancement Using a Mixture of Variational Autoencoders
Mostafa Sadeghi, Xavier Alameda-Pineda
|
👻
Ghosted
|
eess.AS
|
20 |
6 years ago |
| 660 |
Variational Student: Learning Compact and Sparser Networks in Knowledge Distillation Framework
Srinidhi Hegde, Ranjitha Prasad, ... (+2 more)
|
👻
Ghosted
|
cs.LG
|
20 |
6 years ago |
| 661 |
Fast and High-Quality Singing Voice Synthesis System based on Convolutional Neural Networks
Kazuhiro Nakamura, Shinji Takaki, ... (+4 more)
|
👻
Ghosted
|
eess.AS
|
20 |
6 years ago |
| 662 |
Deep clustering with concrete k-means
Boyan Gao, Yongxin Yang, ... (+2 more)
|
👻
Ghosted
|
cs.LG
|
20 |
6 years ago |
| 663 |
H-VECTORS: Utterance-level Speaker Embedding Using A Hierarchical Attention Model
Yanpei Shi, Qiang Huang, Thomas Hain
|
👻
Ghosted
|
cs.CL
|
20 |
6 years ago |
| 664 |
Improve Diverse Text Generation by Self Labeling Conditional Variational Auto Encoder
Yuchi Zhang, Yongliang Wang, ... (+3 more)
|
👻
Ghosted
|
stat.ML
|
20 |
7 years ago |
| 665 |
Automatic assessment of spoken language proficiency of non-native children
Roberto Gretter, Katharina Allgaier, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
20 |
7 years ago |
| 666 |
Clonability of anti-counterfeiting printable graphical codes: a machine learning approach
Olga Taran, Slavi Bonev, Slava Voloshynovskiy
|
👻
Ghosted
|
cs.CR
|
20 |
7 years ago |
| 667 |
Age of Information with Finite Horizon and Partial Updates
David Ramirez, Elza Erkip, H. Vincent Poor
|
👻
Ghosted
|
cs.IT
|
20 |
6 years ago |
| 668 |
Unifying Isolated and Overlapping Audio Event Detection with Multi-Label Multi-Task Convolutional Recurrent Neural Networks
Huy Phan, Oliver Y. Chén, ... (+5 more)
|
👻
Ghosted
|
cs.LG
|
20 |
7 years ago |
| 669 |
Provably Accelerated Randomized Gossip Algorithms
Nicolas Loizou, Michael Rabbat, Peter Richtárik
|
👻
Ghosted
|
math.OC
|
20 |
7 years ago |
| 670 |
Symbol-level precoding is symbol-perturbed ZF when energy Efficiency is sought
Yatao Liu, Wing-Kin Ma
|
👻
Ghosted
|
cs.IT
|
20 |
8 years ago |
| 671 |
Antenna Selection for Large-Scale MIMO Systems with Low-Resolution ADCs
Jinseok Choi, Junmo Sung, ... (+2 more)
|
👻
Ghosted
|
cs.IT
|
20 |
8 years ago |
| 672 |
Convergence analysis of the information matrix in Gaussian belief propagation
Jian Du, Shaodan Ma, ... (+3 more)
|
👻
Ghosted
|
cs.LG
|
20 |
9 years ago |
| 673 |
Accelerated Image Reconstruction for Nonlinear Diffractive Imaging
Yanting Ma, Hassan Mansour, ... (+3 more)
|
👻
Ghosted
|
cs.CV
|
20 |
8 years ago |
| 674 |
Multimodal Signal Processing and Learning Aspects of Human-Robot Interaction for an Assistive Bathing Robot
A. Zlatintsi, I. Rodomagoulakis, ... (+5 more)
|
👻
Ghosted
|
cs.MM
|
20 |
8 years ago |
| 675 |
Character Proposal Network for Robust Text Extraction
Shuye Zhang, Mude Lin, ... (+3 more)
|
👻
Ghosted
|
cs.CV
|
20 |
10 years ago |
| 676 |
An Energy-Efficient Compressive Sensing Framework Incorporating Online Dictionary Learning for Long-term Wireless Health Monitoring
Kai Xu, Yixing Li, Fengbo Ren
|
👻
Ghosted
|
cs.IT
|
20 |
10 years ago |
| 677 |
Unsupervised Spoken Term Detection with Spoken Queries by Multi-level Acoustic Patterns with Varying Model Granularity
Cheng-Tao Chung, Chun-an Chan, Lin-shan Lee
|
👻
Ghosted
|
cs.CL
|
20 |
10 years ago |
| 678 |
Quantum Kernel-Based Long Short-term Memory
Yu-Chao Hsu, Tai-Yu Li, Kuan-Cheng Chen
|
👻
Ghosted
|
quant-ph
|
20 |
1 year ago |
| 679 |
Early Dementia Detection Using Multiple Spontaneous Speech Prompts: The PROCESS Challenge
Fuxiang Tao, Bahman Mirheidari, ... (+13 more)
|
👻
Ghosted
|
cs.SD
|
20 |
1 year ago |
| 680 |
Generating Is Believing: Membership Inference Attacks against Retrieval-Augmented Generation
Yuying Li, Gaoyang Liu, ... (+2 more)
|
👻
Ghosted
|
cs.CR
|
20 |
2 years ago |
| 681 |
Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness
Sicheng Yang, Zunnan Xu, ... (+5 more)
|
💤
Eternal Rest
|
cs.MM
|
20 |
2 years ago |
| 682 |
CHAPTER: Exploiting Convolutional Neural Network Adapters for Self-supervised Speech Models
Zih-Ching Chen, Yu-Shun Sung, Hung-yi Lee
|
👻
Ghosted
|
eess.AS
|
20 |
3 years ago |
| 683 |
Fast and parallel decoding for transducer
Wei Kang, Liyong Guo, ... (+7 more)
|
💀
404 Not Found
|
eess.AS
|
20 |
3 years ago |
| 684 |
Blood Oxygen Saturation Estimation from Facial Video via DC and AC components of Spatio-temporal Map
Yusuke Akamatsu, Yoshifumi Onishi, Hitoshi Imaoka
|
👻
Ghosted
|
eess.IV
|
20 |
3 years ago |
| 685 |
RCDPT: Radar-Camera fusion Dense Prediction Transformer
Chen-Chou Lo, Patrick Vandewalle
|
👻
Ghosted
|
cs.CV
|
20 |
3 years ago |
| 686 |
Massively Multilingual ASR on 70 Languages: Tokenization, Architecture, and Generalization Capabilities
Andros Tjandra, Nayan Singhal, ... (+5 more)
|
👻
Ghosted
|
cs.CL
|
20 |
3 years ago |
| 687 |
Progressive Spatio-Temporal Graph Convolutional Network for Skeleton-Based Human Action Recognition
Negar Heidari, Alexandros Iosifidis
|
👻
Ghosted
|
cs.CV
|
19 |
5 years ago |
| 688 |
Parallel waveform synthesis based on generative adversarial networks with voicing-aware conditional discriminators
Ryuichi Yamamoto, Eunwoo Song, ... (+2 more)
|
👻
Ghosted
|
eess.AS
|
19 |
5 years ago |
| 689 |
Acoustics Based Intent Recognition Using Discovered Phonetic Units for Low Resource Languages
Akshat Gupta, Xinjian Li, ... (+2 more)
|
👻
Ghosted
|
cs.CL
|
19 |
5 years ago |
| 690 |
LANCE: Efficient Low-Precision Quantized Winograd Convolution for Neural Networks Based on Graphics Processing Units
Guangli Li, Lei Liu, ... (+3 more)
|
👻
Ghosted
|
cs.CV
|
19 |
6 years ago |
| 691 |
Improving Voice Separation by Incorporating End-to-end Speech Recognition
Naoya Takahashi, Mayank Kumar Singh, ... (+4 more)
|
👻
Ghosted
|
cs.SD
|
19 |
6 years ago |
| 692 |
Speaker independence of neural vocoders and their effect on parametric resynthesis speech enhancement
Soumi Maiti, Michael I Mandel
|
👻
Ghosted
|
cs.SD
|
19 |
6 years ago |
| 693 |
Conditional Mutual Information Neural Estimator
Sina Molavipour, Germán Bassi, Mikael Skoglund
|
👻
Ghosted
|
cs.IT
|
19 |
6 years ago |
| 694 |
Interpretable Self-Attention Temporal Reasoning for Driving Behavior Understanding
Yi-Chieh Liu, Yung-An Hsieh, ... (+4 more)
|
👻
Ghosted
|
cs.CV
|
19 |
6 years ago |
| 695 |
A Generalization of Principal Component Analysis
Samuele Battaglino, Erdem Koyuncu
|
👻
Ghosted
|
cs.LG
|
19 |
6 years ago |
| 696 |
Channel adversarial training for speaker verification and diarization
Chau Luu, Peter Bell, Steve Renals
|
👻
Ghosted
|
cs.SD
|
19 |
6 years ago |
| 697 |
Tuplemax Loss for Language Identification
Li Wan, Prashant Sridhar, ... (+3 more)
|
👻
Ghosted
|
eess.AS
|
19 |
7 years ago |
| 698 |
Using recurrences in time and frequency within U-net architecture for speech enhancement
Tomasz Grzywalski, Szymon Drgas
|
👻
Ghosted
|
cs.LG
|
19 |
7 years ago |
| 699 |
Adversarial Video Compression Guided by Soft Edge Detection
Sungsoo Kim, Jin Soo Park, ... (+5 more)
|
👻
Ghosted
|
eess.IV
|
19 |
7 years ago |
| 700 |
Multi-Exposure Image Fusion Based on Exposure Compensation
Yuma Kinoshita, Taichi Yoshida, ... (+2 more)
|
👻
Ghosted
|
cs.CV
|
19 |
8 years ago |