← Back to Showcase

Project collection · Ongoing

SpeeD-TB

Speech Datasets and Models for Tibeto-Burman Languages of India

Dataset created under the Speech Datasets and Models for Tibeto-Burman Languages (SpeeD-TB), sponsored under Mission Bhashini by Ministry of Electronics and Information Technology (MEITY), Govt of India. The project aimed to create 1,200 hours of speech dataset and ASR models for 6 underresourced, tribal Tibeto-Burman languages of India speoken in Eastern and North-Eastern parts of India.

Explore the language map →
1,476.8
Hours
6
Languages
345468
Entries
2066
Speakers
Project networkLeadership, partners and funders
5 records

Dataset Development

To build a transcribed speech dataset of approximately 200 hours each in 6 Tibeto-Burman languages - Bodo (mainly spoken in Assam), Meetei (mainly spoken in Manipur), Chokri (mainly spoken in Nagaland), Kok Borok (mainly spoken in Tripura), Nyishi (mainly spoken in Arunachal Pradesh) and Toto (mainly spoken in West Bengal)

80% · In Progress

Phone Set Development

To develop a phone set for each of the languages under study.

100% · Completed

Language Model Development

To build a language model for the languages under consideration here.

100% · Completed

ASR Model Development

To build a baseline ASR system for each of the above languages.

50% · In Progress

Dataset and Model Release

To make the dataset and pre-trained and fine-tuned models publicly available through Bhashini / ULCA and also other platforms and sources including GitHub and other appropriate repositories and server under CC-By 4.0 license (for dataset) and AGPL v3 (for the model).

70% · In Progress