Project collection · Ongoing
SpeeD-TB
Speech Datasets and Models for Tibeto-Burman Languages of India
Dataset created under the Speech Datasets and Models for Tibeto-Burman Languages (SpeeD-TB), sponsored under Mission Bhashini by Ministry of Electronics and Information Technology (MEITY), Govt of India. The project aimed to create 1,200 hours of speech dataset and ASR models for 6 underresourced, tribal Tibeto-Burman languages of India speoken in Eastern and North-Eastern parts of India.
- 1,476.8
- Hours
- 6
- Languages
- 345468
- Entries
- 2066
- Speakers
Chief Investigator and Consortium LeaderProf. Bornini Lahiri
Principle InvestigatorDr. Meiraba Takhellambam
Principle InvestigatorDr. Amalesh Gope
Chief Investigator and Consortium Leader (Former)Dr. Ritesh Kumar
Consortium LeaderIndian Institute of Technology Kharagpur
Consortium PartnerTezpur University
Funders / Funding agenciesMinistry of Electronics and Information Technology, Govt of India
Consortium PartnerManipur University
Technology PartnerUnreal Tece LLP
Host / Lead organisations
1
Consortium partners
5
Tezpur University
Tezpur team was responsible for Chokri and Bodo datasets
Manipur University
Responsible for dataset building of Meetei and Kok Borok and development of their phonesets
Daia Tech Pvt Ltd (Karya)
Responsible for the development of the data collection app and community payments
Partners
1
Unreal Tece LLP
Unreal Tece is a spin-off of the project with the platform being used for data management and transcription as its primary product. They now collaborate in all dsta management and preparation work.
Funders / Funding agencies
1
Ministry of Electronics and Information Technology, Govt of India
The project was funded under Mission Bhashini (National Language Translation Mission)