← Back to Showcase

Project collection · Ongoing

SpeeD-TB

Speech Datasets and Models for Tibeto-Burman Languages of India

Dataset created under the Speech Datasets and Models for Tibeto-Burman Languages (SpeeD-TB), sponsored under Mission Bhashini by Ministry of Electronics and Information Technology (MEITY), Govt of India. The project aimed to create 1,200 hours of speech dataset and ASR models for 6 underresourced, tribal Tibeto-Burman languages of India speoken in Eastern and North-Eastern parts of India.

Explore the language map →
1,476.8
Hours
6
Languages
345468
Entries
2066
Speakers
Project networkLeadership, partners and funders

Languages of SpeeD-TB

Loading map…

Tibeto-Burman

Toto

Toto (ISO 639-3: txo, Glottocode: toto1302) is a member of the Tibeto-Burman language family. Within this family, it is generally classified under the Sal group or grouped alongside Dhimal in the tentative Dhimalish branch, though its precise sub-classification remains subject to ongoing academic debate due to limited historical data.

According to the 2011 Census of India, there are approximately 1,600 members of Toto tribe who identify themseleves as Toto speakers. However, this total community-wide population is distinct from the small, select group of speakers which actually speak the language. The language is currently classified as critically endangered due to the small size of the community and the increasing influence of neighbouring major languages.

Distribution in India

Toto is spoken almost exclusively in Totopara, a small village enclave located in the Alipurduar district (formerly part of the Jalpaiguri district) of West Bengal, India, situated along the border with Bhutan.

Major Grammatical Features

  • Phonology: Toto features a relatively small inventory of vowels and consonants. It is widely considered to have a tonal or register distinction, though the exact phonological behaviour of its pitch system remains under-documented and uncertain.
  • Morphology: The language is predominantly agglutinating. Nouns are inflected for case using postpositions (including ergative/agentive, genitive, accusative, dative, and locative markers). Verbs inflect for tense, aspect, and negation, although the complexity of verbal agreement patterns appears less pronounced than in other Tibeto-Burman languages of the region.
  • Syntax: The basic word order in Toto is Subject-Object-Verb (SOV). It exhibits head-final typological features, with modifiers such as genitives and relative clauses typically preceding the head noun.

Further Reading

Explore relationships → View on language map →

Tibeto-Burman

Nyishi

Nyishi (ISO 639-3: njz, Glottocode: nyis1236) belongs to the Sino-Tibetan (Tibeto-Burman) language family, specifically placed within the Western Tani branch of the post-Tibetan or Tani group.

Speakers

According to the 2011 Census of India, there are approximately 405,100 speakers of Nyishi (enumerated largely under the head "Nissi/Nyishi"). This general speaker population estimate represents the ethnic language community across its traditional geographic range; it is distinct from, and should not be confused with, the small, specific cohort of speakers documented within the localized LiFE corpus.

Distribution in India

The principal region of the Nyishi-speaking population is the Indian state of Arunachal Pradesh, particularly concentrated in the districts of Papum Pare, Kurung Kumey, Kra Daadi, East Kameng, Kamle, and Lower Subansiri. It is also spoken by smaller communities in the adjoining Darrang and Lakhimpur districts of Assam.

Major grammatical features

  • Phonology: Nyishi features a typical Tani vowel system with central vowels (often analysed as containing /ɨ/ and /ə/ alongside /i, e, a, o, u/). Vowel length can be phonemically contrastive. Consonantal contrasts include voiced and voiceless stops, but consonant clusters in syllable-onset position are highly restricted.
  • Morphology: The language is predominantly agglutinative. Nouns are marked for case (including agentive/ergative, accusative, genitive, dative, and locative) using postposed particles. Nyishi possesses a rich numeral classifier system where classifiers are obligatory when counting nouns. Verbs are morphologically complex, inflecting for tense, aspect, mood, and evidentiality via suffixation.
  • Syntax: The basic word order is Subject-Object-Verb (SOV). It is a postpositional language. Within the noun phrase, modifiers such as adjectives, numerals, and demonstratives typically follow the head noun.

Further reading

Explore relationships → View on language map →

Tibeto-Burman

Chokri

Chokri (also known as Chakru, ISO 639-3: nri, Glottocode: chok1243) belongs to the Tibeto-Burman language family, specifically classified within the Angami-Pochuri branch of the Southern Tibeto-Burman subgroup.

Speakers

According to the 2011 Census of India, there are approximately 91,257 speakers of Chakru/Chokri in India.

Distribution in India

Chokri is predominantly spoken in the northeastern state of Nagaland, India. It is concentrated heavily in the Phek district, particularly across the Pfütsero, Chetheba, and Chazouba administrative circles.

Major grammatical features

  • Phonology: Chokri is a tonal language characterised by a complex pitch/tone system that distinguishes lexical meaning. Its consonant inventory is notable for contrasting aspirated and unaspirated stops, alongside a series of voiceless sonorants (including voiceless nasals and laterals) typical of Angami-Pochuri languages.
  • Morphology: The language is predominantly agglutinative. Morphological processes rely heavily on suffixation and prefixation for word formation. Nouns take possessive prefixes corresponding to person and number, and verbs are modified by a range of aspectual, modal, and directional markers.
  • Syntax: Chokri exhibits a basic Subject-Object-Verb (SOV) constituent word order. It is a postpositional language where modifiers like numerals and demonstratives generally follow the head noun, while relative clauses can precede or follow the noun they modify.

Further reading

Explore relationships → View on language map →

Tibeto-Burman

Kok Borok

Kok Borok (ISO 639-3: trp, Glottocode: kokb1239), also known as Kokborok or Tripuri, belongs to the Tibeto-Burman language family. Within this family, it is classified under the Bodo-Garo branch of the Sal subfamily.

Speakers

According to the Census of India 2011, there are approximately 1,011,294 speakers of "Tripuri" (which includes Kok Borok as the dominant variety) in India.

Distribution in India

The principal region of Kok Borok speakers is the state of Tripura in Northeast India, where it holds official status. Smaller communities of speakers are located in the neighbouring states of Assam and Mizoram, as well as in adjacent border areas of Bangladesh.

Major Grammatical Features

  • Phonology: Kok Borok features a distinction between high and low register tones. The vowel inventory includes the back unrounded vowel /ɯ/ (often written as y), typical of Bodo-Garo languages. It has a contrast between aspirated and unaspirated voiceless stops, and syllable-final glottal stops /ʔ/ are common.
  • Morphology: It is predominantly agglutinative. Nouns mark grammatical relations (such as ergative, dative, genitive, and locative) using postposed clitics or suffixes. Verbs undergo complex derivation and inflection for tense, aspect, mood, and polarity via prefixation and suffixation.
  • Syntax: The basic, default word order is Subject-Object-Verb (SOV). The language exhibits an ergative-absolutive alignment system where the agent of a transitive verb is marked with an ergative suffix, while intransitive subjects and transitive objects remain unmarked or receive accusative marking depending on animacy and definiteness.

Further Reading

  • Search for Kok Borok resources on Wikipedia.
  • Detailed classification and structural metadata are available on Glottolog.
  • Standard identification codes can be verified via ISO 639-3.
Explore relationships → View on language map →

Tibeto-Burman

Bodo

Bodo (ISO 639-3: brx, Glottocode: bodo1269) belongs to the Tibeto-Burman language family, specifically belonging to the Bodo-Garo subgroup of the Sal language group.

Speakers

According to the official 2011 Census of India, there are approximately 1.48 million native Bodo speakers.

Distribution in India

The principal concentration of Bodo speakers is in the state of Assam, particularly within the autonomous Bodoland Territorial Region. Smaller speech communities are also distributed across adjacent districts in the states of West Bengal, Meghalaya, and Nagaland.

Major grammatical features

  • Phonology: Bodo is a tonal language, generally described as possessing two distinct lexical tones (high and low). Its consonant inventory features a contrast between aspirated and unaspirated voiceless stops, alongside the prominent use of the voiceless velar fricative /x/.
  • Morphology: The language is predominantly agglutinative. Nouns are marked for grammatical cases—such as nominative, accusative, dative, genitive, locative, and instrumental—via the suffixation of postpositional markers. Verbs exhibit complex inflectional morphology, carrying suffixes to denote tense, aspect, mood, causation, and negation.
  • Syntax: Bodo regularly employs a Subject-Object-Verb (SOV) default word order. It utilises postpositions rather than prepositions. While genitives and demonstratives typically precede the head noun, adjectives and numeral classifiers generally follow it.

Further reading

Explore relationships → View on language map →

Tibeto-Burman

Manipuri

Manipuri (also known as Meitei and Meetei; ISO 639-3: mni, Glottocode: meit1246 / mani1292) belongs to the Tibeto-Burman language family. Its precise subgrouping within Tibeto-Burman remains a subject of academic debate, often classified within its own independent branch or grouped tentatively with the Kuki-Chin-Naga languages.

Speakers

According to the 2011 Census of India, there are approximately 1.76 million native speakers of Manipuri in India.

Distribution in India

The language is primarily spoken in the northeastern state of Manipur, where it serves as the official state language and the lingua franca among diverse ethnic groups. Significant speaker communities also exist in the neighbouring states of Assam, Tripura, and Nagaland.

Major Grammatical Features

  • Phonology: Manipuri is a tonal language, traditionally analysed as having two contrastive tones (high and low/level). The phoneme inventory consists of approximately 28 segmental phonemes, including six vowels and a distinction between aspirated and unaspirated stops.
  • Morphology: Highly agglutinative. Grammatical relations are marked predominantly through suffixation, though prefixation plays a crucial role in nominal derivation and pronominal possession. There is no grammatical gender; instead, natural gender is indicated lexically or through specific markers.
  • Syntax: The basic word order is Subject-Object-Verb (SOV). It is a split-ergative or agentive-aligning language where case markers (such as the agentive/nominative and accusative enclitics) are applied pragmatically depending on emphasis, definiteness, and control.

Further reading

Explore relationships → View on language map →