Phonetically-Augmented Discriminative Rescoring for Voice Search Error Correction
By: Christophe Van Gysel , Maggie Wu , Lyan Verwimp and more
Potential Business Impact:
Helps voice search understand movie titles better.
End-to-end (E2E) Automatic Speech Recognition (ASR) models are trained using paired audio-text samples that are expensive to obtain, since high-quality ground-truth data requires human annotators. Voice search applications, such as digital media players, leverage ASR to allow users to search by voice as opposed to an on-screen keyboard. However, recent or infrequent movie titles may not be sufficiently represented in the E2E ASR system's training data, and hence, may suffer poor recognition. In this paper, we propose a phonetic correction system that consists of (a) a phonetic search based on the ASR model's output that generates phonetic alternatives that may not be considered by the E2E system, and (b) a rescorer component that combines the ASR model recognition and the phonetic alternatives, and select a final system output. We find that our approach improves word error rate between 4.4 and 7.6% relative on benchmarks of popular movie titles over a series of competitive baselines.
Similar Papers
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
Computation and Language
Helps computers understand many people talking at once.
ASR Error Correction in Low-Resource Burmese with Alignment-Enhanced Transformers using Phonetic Features
Computation and Language
Fixes mistakes in spoken Burmese words.
Evaluating ASR Confidence Scores for Automated Error Detection in User-Assisted Correction Interfaces
Human-Computer Interaction
Makes computer speech-to-text less helpful.