← Back to Showcase

Project collection · Ongoing

SpeeD-TB

Speech Datasets and Models for Tibeto-Burman Languages of India

Dataset created under the Speech Datasets and Models for Tibeto-Burman Languages (SpeeD-TB), sponsored under Mission Bhashini by Ministry of Electronics and Information Technology (MEITY), Govt of India. The project aimed to create 1,200 hours of speech dataset and ASR models for 6 underresourced, tribal Tibeto-Burman languages of India speoken in Eastern and North-Eastern parts of India.

Explore the language map →
1,476.8
Hours
6
Languages
345468
Entries
2066
Speakers
Project networkLeadership, partners and funders
Rationale

Over the last decade or so, research in speech technologies has seen a rapid and successful shift towards exclusively data-driven techniques such as machine learning and deep learning methods. Over the years, experiments with well-resourced languages such as English have demonstrated the success of these systems given sufficient data for training the systems. However, barring a handful of languages, this technological revolution has escaped most of the languages (including the officially supported, scheduled languages) spoken in India. This could be gauged from the commercial support for very few Indian languages across different speech-based products - Amazon Alexa supports Hindi among seven other international languages; Google Home supports 13 languages, including Hindi, as the only Indian language; Microsoft supports Indian English, Hindi, Tamil, Telugu, Gujarati, and Marathi for its ASR systems - there is no support whatsoever for most of the other Indian languages, especially languages belonging to the Tibeto-Burman and Austro-Asiatic language families. One of the main reasons behind this is the non-availability of sufficient speech datasets for most of the Indian languages. This is even more so for the languages belonging to the Tibeto-Burman and Austro-Asiatic language families, largely spoken in Eastern and North-Eastern parts of India. Out of the 22 scheduled languages, Bodo and Meetei belong to the Tibeto-Burman language family and Santhali belong to the Munda sub-group of Austro-Asiatic language family. As per our survey of the resources and corpora available for building speech technologies in Indian languages, the resources available for these three languages are listed below - * Approximately 177 hours of speech data collected from 456 speakers is available in Bodo - this dataset is provided by the LDC-IL Speech Corpus. * Slightly over 156 hours of speech data collected from 620 speakers is available in Meetei through the LDC-IL Speech Corpus

Clearly, these datasets are not sufficient for the modern data-hungry deep learning systems which are typically trained on much larger datasets of thousands of hours. Furthermore, some of the other major and state official languages do not have even these minimal resources.

This project is envisioned to alleviate this situation with respect to languages under study in the project. In the following subsections, we discuss the different parts and aspects, proposed methodology and expected results of the project.

Objectives

In this initial phase, the main objective of the project is to build a speech dataset of at least 1,200 hours consisting of around 200 hours in 6 Indian languages from the Tibeto-Burman language family. The project will also prepare phone sets, language models and baseline models for speech recognition in these languages. In the next stages / phases of the project, we plan to expand to more languages including those from Austro-Asiatic and Tibeto-Burman language families.