Turbocharge Speech Understanding with Pilot Inference

November 22, 2023 · Declared Dead · + Add venue

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Rongxiang Wang, Felix Xiaozhu Lin arXiv ID 2311.17065 Category eess.AS: Audio & Speech Cross-listed cs.CL, cs.LG Citations 4 Last Checked 3 months ago

Abstract

Modern speech understanding (SU) runs a sophisticated pipeline: ingesting streaming voice input, the pipeline executes encoder-decoder based deep neural networks repeatedly; by doing so, the pipeline generates tentative outputs (called hypotheses), and periodically scores the hypotheses. This paper sets to accelerate SU on resource-constrained edge devices. It takes a hybrid approach: to speed up on-device execution; to offload inputs that are beyond the device's capacity. While the approach is well-known, we address SU's unique challenges with novel techniques: (1) late contextualization, which executes a model's attentive encoder in parallel to the input ingestion; (2) pilot inference, which mitigates the SU pipeline's temporal load imbalance; (3) autoregression offramps, which evaluate offloading decisions based on pilot inferences and hypotheses. Our techniques are compatible with existing speech models, pipelines, and frameworks; they can be applied independently or in combination. Our prototype, called PASU, is tested on Arm platforms with 6 - 8 cores: it delivers SOTA accuracy; it reduces the end-to-end latency by 2x and reduces the offloading needs by 2x.

📄 View on arXiv 🌐 View on ar5iv 📑 PDF 🎉 Report Code Found

Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

📜 Similar Papers

In the same crypt — Audio & Speech

R.I.P. 👻 Ghosted

ASVspoof 2019: Future Horizons in Spoofed and Fake Audio Detection

Massimiliano Todisco, Xin Wang, ... (+8 more)

eess.AS 🏛 Interspeech 📚 736 cites 7 years ago

R.I.P. 👻 Ghosted

LPCNet: Improving Neural Speech Synthesis Through Linear Prediction

Jean-Marc Valin, Jan Skoglund

eess.AS 🏛 ICASSP 📚 489 cites 7 years ago

R.I.P. 👻 Ghosted

VoiceFilter: Targeted Voice Separation by Speaker-Conditioned Spectrogram Masking

Quan Wang, Hannah Muckenhirn, ... (+8 more)

eess.AS 🏛 Interspeech 📚 413 cites 7 years ago

R.I.P. 👻 Ghosted

TERA: Self-Supervised Learning of Transformer Encoder Representation for Speech

Andy T. Liu, Shang-Wen Li, Hung-yi Lee

eess.AS 🏛 IEEE/ACM TASLP 📚 397 cites 5 years ago

R.I.P. 👻 Ghosted

Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders

Andy T. Liu, Shu-wen Yang, ... (+3 more)

eess.AS 🏛 ICASSP 📚 393 cites 6 years ago

R.I.P. 👻 Ghosted

Utterance-level Aggregation For Speaker Recognition In The Wild

Weidi Xie, Arsha Nagrani, ... (+2 more)

eess.AS 🏛 ICASSP 📚 365 cites 7 years ago

Died the same way — 👻 Ghosted

R.I.P. 👻 Ghosted

Federated Learning: Strategies for Improving Communication Efficiency

Jakub Konečný, H. Brendan McMahan, ... (+4 more)

cs.LG 🏛 arXiv 📚 5.2K cites 9 years ago

R.I.P. 👻 Ghosted

In-Datacenter Performance Analysis of a Tensor Processing Unit

Norman P. Jouppi, Cliff Young, ... (+73 more)

cs.AR 🏛 ISCA 📚 5.1K cites 9 years ago

R.I.P. 👻 Ghosted

Deep Convolutional Neural Networks for Computer-Aided Detection: CNN Architectures, Dataset Characteristics and Transfer Learning

Hoo-Chang Shin, Holger R. Roth, ... (+7 more)

cs.CV 🏛 IEEE TMI 📚 4.9K cites 10 years ago

R.I.P. 👻 Ghosted

Explanation in Artificial Intelligence: Insights from the Social Sciences

Tim Miller

cs.AI 🏛 AI 📚 4.9K cites 8 years ago