Score: 0

Edge-Based Speech Transcription and Synthesis for Kinyarwanda and Swahili Languages

Published: October 18, 2025 | arXiv ID: 2510.16497v1

By: Pacome Simon Mbonimpa , Diane Tuyizere , Azizuddin Ahmed Biyabani and more

Potential Business Impact:

Lets phones understand and speak Kinyarwanda and Swahili.

Business Areas:

Speech Recognition Data and Analytics, Software

This paper presents a novel framework for speech transcription and synthesis, leveraging edge-cloud parallelism to enhance processing speed and accessibility for Kinyarwanda and Swahili speakers. It addresses the scarcity of powerful language processing tools for these widely spoken languages in East African countries with limited technological infrastructure. The framework utilizes the Whisper and SpeechT5 pre-trained models to enable speech-to-text (STT) and text-to-speech (TTS) translation. The architecture uses a cascading mechanism that distributes the model inference workload between the edge device and the cloud, thereby reducing latency and resource usage, benefiting both ends. On the edge device, our approach achieves a memory usage compression of 9.5% for the SpeechT5 model and 14% for the Whisper model, with a maximum memory usage of 149 MB. Experimental results indicate that on a 1.7 GHz CPU edge device with a 1 MB/s network bandwidth, the system can process a 270-character text in less than a minute for both speech-to-text and text-to-speech transcription. Using real-world survey data from Kenya, it is shown that the cascaded edge-cloud architecture proposed could easily serve as an excellent platform for STT and TTS transcription with good accuracy and response time.

TeluguST-46: A Benchmark Corpus and Comprehensive Evaluation for Telugu-English Speech Translation

Computation and Language

Translates Telugu speech to English better.

8 Dec 2025 1

87%

Democratizing Agentic AI with Fast Test-Time Scaling on the Edge

Machine Learning (CS)

Lets small computers think like big ones.

29 Aug 2025 2

87%

End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs

Computation and Language

Lets computers translate spoken words to another language.

11 Oct 2025 1

View PDF Login to Bookmark

Country of Origin

🇺🇸 United States

Page Count

11 pages

Edge-Based Speech Transcription and Synthesis for Kinyarwanda and Swahili Languages

Lets phones understand and speak Kinyarwanda and Swahili.

Technical Abstract

TeluguST-46: A Benchmark Corpus and Comprehensive Evaluation for Telugu-English Speech Translation

Democratizing Agentic AI with Fast Test-Time Scaling on the Edge

End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs