Document Type
Conference Proceeding
Publication Title
Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
Abstract
This paper presents FOOCTTS, an automatic pipeline for a football commentator that generates speech with background crowd noise. The application gets the text from the user, applies text pre-processing such as vowelization, followed by the commentator's speech synthesizer. Our pipeline included Arabic automatic speech recognition for data labeling, CTC segmentation, transcription vowelization to match speech, and fine-tuning the TTS. Our system is capable of generating speech with its acoustic environment within limited 15 minutes of football commentator recording. Our prototype is generalizable and can be easily applied to different domains and languages.
First Page
5249
Last Page
5250
Publication Date
8-2023
Keywords
speech recognition, text-to-speech
Recommended Citation
M. Baali and A. Ali, "FOOCTTS: Generating Arabic Speech with Acoustic Environment for Football Commentator," Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, vol. 2023-August, pp. 5249 - 5250, Aug 2023.
Additional Links
ISCA Archive link: https://www.isca-archive.org/interspeech_2023/baali23b_interspeech.html
Comments
IR conditions: non-described