Bringing NURC/SP to Digital Life: the Role of Open-source Automatic Speech Recognition Models

October 14, 2022 ยท Declared Dead ยท ๐Ÿ› Anais do XIX Encontro Nacional de Inteligรชncia Artificial e Computacional (ENIAC 2022)

๐Ÿ‘ป CAUSE OF DEATH: Ghosted
No code link whatsoever

"No code URL or promise found in abstract"

Evidence collected by the PWNC Scanner

Authors Lucas Rafael Stefanel Gris, Arnaldo Candido Junior, Vinรญcius G. dos Santos, Bruno A. Papa Dias, Marli Quadros Leite, Flaviane Romani Fernandes Svartman, Sandra Aluรญsio arXiv ID 2210.07852 Category cs.CL: Computation & Language Cross-listed cs.SD, eess.AS Citations 3 Venue Anais do XIX Encontro Nacional de Inteligรชncia Artificial e Computacional (ENIAC 2022) Last Checked 5 months ago
Abstract
The NURC Project that started in 1969 to study the cultured linguistic urban norm spoken in five Brazilian capitals, was responsible for compiling a large corpus for each capital. The digitized NURC/SP comprises 375 inquiries in 334 hours of recordings taken in Sรฃo Paulo capital. Although 47 inquiries have transcripts, there was no alignment between the audio-transcription, and 328 inquiries were not transcribed. This article presents an evaluation and error analysis of three automatic speech recognition models trained with spontaneous speech in Portuguese and one model trained with prepared speech. The evaluation allowed us to choose the best model, using WER and CER metrics, in a manually aligned sample of NURC/SP, to automatically transcribe 284 hours.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

๐Ÿ“œ Similar Papers

In the same crypt โ€” Computation & Language

๐ŸŒ… ๐ŸŒ… Old Age

Attention Is All You Need

Ashish Vaswani, Noam Shazeer, ... (+6 more)

cs.CL ๐Ÿ› NeurIPS ๐Ÿ“š 166.0K cites 9 years ago

Died the same way โ€” ๐Ÿ‘ป Ghosted