← Back to Showcase

Project collection · Ongoing

Project BoLI

Reclaiming Data Sovereignty. Revolutionising AI Development for 1,000+ Indian Languages

Project BoLI, the flagship initiative of UnReaL-TecE LLP, is building high-quality datasets to train and benchmark next-generation AI tasks—from speech-to-text to advanced grammatical reasoning—across more than 1,000 underserved Indian languages and dialects. We are pioneering a globally unique data governance framework where all released datasets, derivative assets, and downstream models are permanently co-owned by the contributing communities. Under our groundbreaking BoLI License, commercial data usage is strictly treated as a revocable license rather than a transfer of ownership, ensuring total accountability in the era of LLMs.

Explore the language map →
268.3
Hours
64
Languages
315163
Entries
170
Speakers

Explore the collection

See how Project BoLI connects

Follow the relationships between projects, people, languages, questionnaires, resources, and public data.

Loading network…

Explore full network
Project networkLeadership, partners and funders

Languages of BoLI

Loading map…

Tibeto-Burman

Toto

Toto (ISO 639-3: txo, Glottocode: toto1302) is a member of the Tibeto-Burman language family. Within this family, it is generally classified under the Sal group or grouped alongside Dhimal in the tentative Dhimalish branch, though its precise sub-classification remains subject to ongoing academic debate due to limited historical data.

According to the 2011 Census of India, there are approximately 1,600 members of Toto tribe who identify themseleves as Toto speakers. However, this total community-wide population is distinct from the small, select group of speakers which actually speak the language. The language is currently classified as critically endangered due to the small size of the community and the increasing influence of neighbouring major languages.

Distribution in India

Toto is spoken almost exclusively in Totopara, a small village enclave located in the Alipurduar district (formerly part of the Jalpaiguri district) of West Bengal, India, situated along the border with Bhutan.

Major Grammatical Features

  • Phonology: Toto features a relatively small inventory of vowels and consonants. It is widely considered to have a tonal or register distinction, though the exact phonological behaviour of its pitch system remains under-documented and uncertain.
  • Morphology: The language is predominantly agglutinating. Nouns are inflected for case using postpositions (including ergative/agentive, genitive, accusative, dative, and locative markers). Verbs inflect for tense, aspect, and negation, although the complexity of verbal agreement patterns appears less pronounced than in other Tibeto-Burman languages of the region.
  • Syntax: The basic word order in Toto is Subject-Object-Verb (SOV). It exhibits head-final typological features, with modifiers such as genitives and relative clauses typically preceding the head noun.

Further Reading

Explore relationships → View on language map →
Projects in this collectionBoLI Toto TranslationBoLI Toto Narration

Tibeto-Burman

Pangal

Pangal (Glottolog: pang1284) is a Tibeto-Burman language. It is highly similar to and classified under the same ISO 639-3 code (mni) as Meitei (Manipuri), and is widely analyzed as a distinct sociolect or dialect of Meitei spoken by the Manipuri Muslim (Pangal) community.

Speakers

The broader Pangal community in Northeast India is estimated to comprise approximately 240,000 speakers according to the 2011 Census of India, based on the reported Muslim population of Manipur who list Meitei/Manipuri as their mother tongue.

Distribution in India

Pangal speakers are primarily concentrated in the Imphal Valley of the state of Manipur, particularly within the Thoubal, Imphal West, Imphal East, and Bishnupur districts. Smaller diaspora communities are also found in neighboring states, including Assam and Tripura.

Major grammatical features

  • Phonology: Structurally shares the phonological framework of Meitei, utilizing a pitch-accent or two-tone system (distinguishing high and low tones). It features a distinct lexical overlay, incorporating Arabic, Persian, and Urdu loanwords to accommodate Islamic cultural and religious terminology, which introduces non-native phonetic elements in specific registers.
  • Morphology: Highly agglutinative. Grammatical relationships, verb tenses, aspects, and modal distinctions are primarily indicated through extensive prefixation and suffixation. Nominal morphology relies on postpositional clitics to denote case relations.
  • Syntax: Adheres strictly to a Subject-Object-Verb (SOV) default constituent order. It is a pro-drop language, allowing the omission of pronouns when the agent or patient can be inferred from context. Clause chaining using non-finite verbal endings is common.

Further reading

Explore relationships → View on language map →
Projects in this collectionBoLI Pangal Translation

Dravidian

Tulu

Tulu belongs to the Dravidian language family, specifically classified within the South Dravidian branch.

Speakers

According to the 2011 Census of India, there are approximately 1,846,427 native Tulu speakers in India.

Distribution in India

Tulu is principally spoken in the southwestern coastal region of India, a belt traditionally referred to as Tulu Nadu. This region primarily encompasses the Dakshina Kannada and Udupi districts in the state of Karnataka, as well as the northern part of the Kasaragod district in Kerala.

Major grammatical features

  • Phonology: Tulu possesses a typical Dravidian phonemic inventory, including a series of retroflex consonants (/ʈ/, /ɖ/, /ɳ/, /ɭ/). It is highly notable for its vowel system, which includes the close back unrounded vowel (/ɯ/) in addition to the close back rounded vowel (/u/).
  • Morphology: Highly agglutinative. Nouns are marked for number (singular and plural) and decline for several cases, including nominative, accusative, dative, genitive, locative, and instrumental/sociative. Tulu distinguishes three grammatical genders in the singular (masculine, feminine, and neuter), though the masculine and feminine collapse into a common "rational" class in the plural. Verbs conjugate for tense (past, present, future), mood, and person-number-gender agreement.
  • Syntax: Tulu strictly adheres to a Subject-Object-Verb (SOV) default word order. It is a left-branching language, meaning modifiers and relative clauses typically precede their head nouns, and postpositions are used instead of prepositions.

Further reading

Explore relationships → View on language map →
Projects in this collectionBoLI Tulu NarrationBoLI Tulu Translation

Tibeto-Burman

Garo

Garo belongs to the Tibeto-Burman language family, situated within the Bodo-Garo branch of the Sal subfamily.

Speakers

According to the 2011 Census of India, there are approximately 1,145,323 native Garo speakers in India.

Distribution in India

The principal concentration of Garo speakers is in the state of Meghalaya, particularly within the East, West, North, South, and Southwest Garo Hills districts. Significant communities also reside in adjacent areas of Assam (including the Goalpara, Kamrup, and Karbi Anglong districts), Tripura, and West Bengal.

Major grammatical features

  • Phonology: Garo has a relatively simple vowel system (/a, e, i, o, u/) and a consonant inventory characterized by a distinction between aspirated and unaspirated voiceless stops. A defining feature is the glottal stop /ʔ/ (often orthographically represented as q or an apostrophe), which frequently appears in syllable-coda positions. Unlike many Tibeto-Burman languages, Garo is generally considered non-tonal.
  • Morphology: It is highly agglutinative, relying heavily on suffixation. Noun phrases employ a rich system of numeral classifiers (e.g., sak for humans, gong for flat objects) that must accompany numerals. Nouns are marked for case (nominative, accusative, genitive, dative, locative, and instrumental) via postpositional clitics. Verbs are morphologically complex, inflecting for tense, aspect, mood, and negation through chain-like suffixation.
  • Syntax: Garo exhibits a rigid Subject-Object-Verb (SOV) basic word order. It is a postpositional language where modifiers, such as adjectives, typically follow the noun they modify, while genitives and numeral classifiers precede the noun head.

Further reading

Explore relationships → View on language map →

Austro-Asiatic

Khasi

Khasi belongs to the Austroasiatic language family, specifically classified within the Khasic branch.

Speakers

According to the 2011 Census of India, there are approximately 1,431,344 speakers of Khasi in India.

Distribution in India

The language is primarily spoken in the state of Meghalaya, particularly within the East Khasi Hills, West Khasi Hills, South West Khasi Hills, Eastern West Khasi Hills, and Ri-Bhoi districts. Significant speaker communities are also found in the neighbouring state of Assam.

Major Grammatical Features

  • Phonology: Khasi is a non-tonal language. It features a moderately large inventory of consonants, including aspirated stops, and allows complex initial consonant clusters. It contrasts short and long vowels.
  • Morphology: Khasi is largely analytical, though it employs a rich set of derivational prefixes (such as the causative pyn-) and infixes. Nouns are obligatorily marked for grammatical gender using four proclitic articles: u (masculine), ka (feminine), i (diminutive/affectionate), and ki (plural).
  • Syntax: The basic word order is strictly Subject-Verb-Object (SVO). It is a head-initial language, meaning that prepositions are used instead of postpositions, and nouns typically precede their modifying adjectives and genitive phrases.

Further reading

Explore relationships → View on language map →

Austro-Asiatic

Ho

Ho belongs to the Austro-Asiatic language family, specifically situated within the North Munda branch of the Munda sub-family.

Speakers

According to the 2011 Census of India, there are approximately 1.4 million speakers of the Ho language.

Distribution in India

The language is primarily spoken in the state of Jharkhand (especially concentrated in the West Singhbhum and East Singhbhum districts). It is also widely spoken in the adjoining districts of northern Odisha (such as Mayurbhanj, Keonjhar, and Sundargarh) and parts of West Bengal.

Major grammatical features

  • Phonology: The phonemic inventory of Ho includes a distinction between aspirated and unaspirated stops, retroflex consonants, and characteristic Austro-Asiatic glottalized or checked stops (often occurring syllable-finally). Vowel nasalization and length can also carry contrastive functional weight.
  • Morphology: Ho is a highly agglutinative language. Nouns are marked for three grammatical numbers (singular, dual, and plural) and inflected for grammatical case. Verbs are highly complex, incorporating pronominal markers for both the subject and object, alongside suffixes indicating tense, aspect, and mood.
  • Syntax: The default word order is Subject-Object-Verb (SOV). Ho is head-final, utilizing postpositions rather than prepositions, with modifiers generally preceding the nouns they qualify.

Further reading

Explore relationships → View on language map →
Projects in this collectionBoLI Ho Translation

Dravidian

Jenu Kurumba

Jenu Kurumba (ISO 639-3: xuj, Glottocode: jenn1240) belongs to the Southern Dravidian branch of the Dravidian language family. It is closely related to Kannada and other Southern Dravidian varieties spoken in the Nilgiri hills region.

Speakers

Scholarly and diagnostic databases estimate the Jenu Kurumba speaker population to be approximately 35,000 (Endangered Languages Project), though precise census figures are difficult to isolate due to the historical grouping of various Kurumba communities under broader linguistic identities in national surveys.

Distribution in India

The language is primarily spoken in the forested border regions of Karnataka, Tamil Nadu, and Kerala. The core speaker communities reside in the Mysore (Chamarajanagar) and Kodagu districts of Karnataka, as well as adjacent parts of the Nilgiris district in Tamil Nadu and Wayanad district in Kerala.

Major grammatical features

  • Phonology: Jenu Kurumba features a typical Dravidian vowel inventory with five basic vowels distinguishes by length (short and long). The consonant inventory includes characteristic retroflex consonants (such as /ʈ/, /ɖ/, and /ɳ/) and a distinction between voiced and voiceless stops, though voicing may be conditioned by phonetic environment.
  • Morphology: The language is highly agglutinative. Nouns are inflected for case (including nominative, accusative, dative, genitive, locative, and instrumental) and number. It generally maintains a Dravidian gender-marking system (often contrasting rational vs. irrational or masculine/feminine/neuter). Verbs inflect for tense (past, non-past), aspect, mood, and person-number-gender agreement.
  • Syntax: Word order is predominantly Subject-Object-Verb (SOV). Jenu Kurumba is head-final, utilizing postpositions rather than prepositions, and relative clauses generally precede the noun they modify.

Further reading

Explore relationships → View on language map →

Dravidian

Kodava

Kodava (also known as Kodava Takk) belongs to the Dravidian language family, specifically classified under the Southern Dravidian (South Dravidian I) subgroup. It is closely related to Tamil, Malayalam, and Kannada.

Speakers

According to the 2011 Census of India, there are approximately 113,857 speakers of Kodagu/Coorgi in India.

Distribution in India

The language is primarily spoken in the Kodagu (Coorg) district in the southwestern region of the state of Karnataka, India. Due to migration, communities of speakers are also found in neighboring districts of Karnataka and Kerala, as well as in major Indian metropolitan areas.

Major grammatical features

  • Phonology: Kodava is highly distinctive among Dravidian languages for its vowel system, which features a set of high central unrounded vowels (/ɯ/, /ɯː/) and mid-central vowels (/ë/, /ëː/), alongside the typical Dravidian five-vowel system. It maintains a contrast between short and long vowels.
  • Morphology: It is an agglutinative language. Nouns are marked for cases (including nominative, accusative, genitive, dative, locative, instrumental, and sociative) and number. The gender system distinguishes masculine, feminine, and neuter, though plural forms often neutralize some gender distinctions. Verbs are inflected for tense (past, present, future), aspect, mood, and show agreement with the subject in person, number, and gender.
  • Syntax: The basic word order is Subject-Object-Verb (SOV). Being a head-final language, it exclusively uses postpositions rather than prepositions, and modifiers (such as adjectives and relative clauses) precede the nouns they modify.

Further reading

Explore relationships → View on language map →
Projects in this collectionBoLI Kodava Translation

Dravidian

Erava

Erava (often identified with or closely related to Ravula) belongs to the Dravidian language family, specifically classified within the Southern Dravidian subgroup.

Speakers

The total number of Erava speakers remains highly uncertain due to varying classifications under the "Yerava" or "Ravula" labels. However, scholarly estimates and tribal surveys place the ethnic speaker population at approximately 25,000 to 30,000, as indexed in reports like the Census of India 2011.

Distribution in India

Erava is primarily spoken in the Southern region of India. Its principal concentration is in the Kodagu (Coorg) district of Karnataka, with some speakers also residing in the neighboring districts of Kerala.

Major grammatical features

  • Phonology: Erava features a typical South Dravidian phonological system, including a distinction between short and long vowels, and a consonant inventory that features retroflex stops (/ʈ/, /ɖ/), nasals (/ɳ/), and laterals (/ɭ/).
  • Morphology: It is a highly agglutinative language. Nouns inflect for case (including nominative, accusative, dative, genitive, and locative) and number. Verbs inflect for tense (distinguishing past and non-past), mood, and aspect, typically agreeing with the subject in person, number, and gender.
  • Syntax: The basic word order is Subject-Object-Verb (SOV). It is a head-final language employing postpositions rather than prepositions, and relative clauses precede the nouns they modify.

Further reading

For more information, see the Wikipedia Search for Erava Language.

Explore relationships → View on language map →
Projects in this collectionBoLI Erava Narration

Dravidian

Kurukh

Kurukh (ISO 639-3: kru, Glottocode: kuru1301) belongs to the Dravidian language family, specifically classified within the North Dravidian subgroup. It shares close genetic and structural relationships with Malto and Brahui.

Speakers

According to the 2011 Census of India, Kurukh (returned largely under the name 'Kurukh/Oraon') has approximately 1,988,350 speakers.

Distribution in India

The principal concentration of Kurukh speakers is located in the Chota Nagpur Plateau region of eastern India. The language is spoken primarily in the following states: * Jharkhand * Chhattisgarh * Odisha * West Bengal * Bihar * Minor populations of speakers also exist in parts of Assam and Tripura, often associated with historical migrations to tea garden communities.

Major grammatical features

  • Phonology: Kurukh features a standard Dravidian five-vowel system (both short and long varieties: /a/, /e/, /i/, /o/, /u/). The consonant inventory includes a distinctive contrast between dental and retroflex stops, as well as a glottal fricative (/h/) and, in some dialects, a glottal stop.
  • Morphology: It is a highly agglutinating and suffixing language. Nouns are inflected for number (singular and plural) and eight grammatical cases (nominative, accusative, genitive, dative, instrumental, ablative, locative, and vocative). The gender system is structurally unique: it distinguishes masculine (male humans) from non-masculine (females, animals, and inanimate objects). Verbs conjugate extensively for tense, aspect, mood, and person-number-gender agreement with the subject.
  • Syntax: Syntactic structure is consistently Subject-Object-Verb (SOV). Modifiers, such as genitives and adjectives, strictly precede the nouns they modify.

Further reading

Explore relationships → View on language map →

Dravidian

Markodi

Markodi is classified as a member of the Dravidian language family. It belongs to the South Dravidian subgroup, showing genetic or areal relationships with major regional languages such as Malayalam, Tulu, or Kannada.

Speakers

The total speaker population of Markodi is currently unknown, as it is not separately enumerated in the Census of India or other major demographic databases.

Distribution in India

Markodi is located in the Kasaragod district of Kerala, India, close to the border of Karnataka (approximate coordinates: 12.51° N, 74.99° E). This borderland region is highly multilingual, characterized by intense contact between Malayalam, Kannada, Tulu, and various localized tribal or minority lects.

Major Grammatical Features

Due to the lack of dedicated grammatical descriptions, specific linguistic features of Markodi remain unconfirmed but are expected to align with the typological profile of South Dravidian languages: * Phonology: Likely features a contrast between short and long vowels, a rich inventory of Dravidian retroflex consonants (/ʈ/, /ɖ/, /ɳ/, /ɭ/), and a lack of initial consonant clusters. The use of Malayalam script in transcriptions suggests phonological alignment with regional Malayalam or Tulu varieties. * Morphology: Highly agglutinating. Nouns are marked for grammatical number (singular/plural) and case (including nominative, accusative, dative, genitive, locative, and sociative) using postpositions or suffixes. Verbs typically inflect for tense (past, present, future), aspect, mood, and person-number-gender (PNG) agreement. * Syntax: Strictly head-final with a default Subject-Object-Verb (SOV) word order, extensive use of relative participles instead of relative clauses, and postpositional phrases.

Further reading

Explore relationships → View on language map →

Indo-Aryan

Siddi Jananga

Siddi Jananga is classified as an Indo-Aryan language variety. Synthesized within a multilingual contact environment, it is closely related to Konkani and Marathi but it has undergone profound structural convergence with Dravidian languages.

Speakers

The exact number of mother-tongue speakers of Siddi Jananga remains uncertain because it is not separately tabulated in official Indian national censuses. However, the ethnic Siddi population in Karnataka is estimated to be between 10,000 and 25,000 individuals according to scholarly and demographic studies.

Distribution in India

The speech community is primarily settled in the state of Karnataka, with major concentrations in the rural and forested taluks of the Uttara Kannada, Belagavi, and Dharwad districts.

Major grammatical features

  • Phonology: The sound system exhibits a typical Indo-Aryan contrast between aspirated and unaspirated stops, alongside a prominent series of retroflex consonants (/ʈ/, /ɖ/, /ɳ/). It also features vowel nasalization and phonetic simplification influenced by the surrounding Kannada-speaking majority.
  • Morphology: Nouns are marked for case using a combination of Indo-Aryan inflections and Dravidian-influenced postpositions. The grammatical gender system shows significant restructuring and simplification compared to standard Konkani, frequently aligning with Kannada's rational/irrational gender divides.
  • Syntax: The syntax is characterized by a stable Subject-Object-Verb (SOV) word order. It makes extensive use of non-finite verb forms, such as conjunctive participles and converbs for clause-chaining, and employs the dative-subject construction to denote physical or psychological experiences.

Further reading

For additional cultural and linguistic context, search the Siddi Jananga language on Wikipedia.

Explore relationships → View on language map →
Projects in this collectionBoLI Siddi Jananga Translation