← Back to Showcase

Project collection · Ongoing

Project BoLI

Reclaiming Data Sovereignty. Revolutionising AI Development for 1,000+ Indian Languages

Project BoLI, the flagship initiative of UnReaL-TecE LLP, is building high-quality datasets to train and benchmark next-generation AI tasks—from speech-to-text to advanced grammatical reasoning—across more than 1,000 underserved Indian languages and dialects. We are pioneering a globally unique data governance framework where all released datasets, derivative assets, and downstream models are permanently co-owned by the contributing communities. Under our groundbreaking BoLI License, commercial data usage is strictly treated as a revocable license rather than a transfer of ownership, ensuring total accountability in the era of LLMs.

Explore the language map →
268.3
Hours
64
Languages
315163
Entries
170
Speakers
Project networkLeadership, partners and funders
73 records

dataset

BoLI Pangal Translation

The Dataset

The current data preview of Pangal is being released as part of the Project BoLI. This preview is a reflection of the full dataset and consists of the following -

  1. Speech Recordings of 200 sentences in the language.
  2. Transcriptions in IPA and Bangali, English
  3. Translations in English (which also act as prompts for the translation sentences)
  4. Detailed speaker metadata, including their demographic, educational and linguistic profile.
  5. Prompt in English and Hindi.

The full dataset contains the following -

  1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones.
  2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks.
  3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations.
  4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc.
  5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors.

Dataset Preparation

The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language.

The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the guidelines for the project using MATra Lab, a part of the LiFE Suite Ecosystem, developed by Unreal Tece LLP.

Project BoLI Guidelines

About Pangal

Pangal is a Tibeto-Burman language, spoken primarily in Imphal East, Manipur by 3,58,000 speakers as per Census 2011 / other sources. The current dataset is primarily recorded by speakers from New Checkon, Imphal East, Manipur. Phonetically, Pangal features a basic six-vowel system: /i/, /e/, /a/, /u/, /o/, and /ə/. The vowel phonemes are categorised by four levels of height (high, mid-high, mid, and low) and three levels of backness (front, central, and back). The language also has six distinct diphthongs - /əi/, /əu/, /ai/, /oi/, /au/, /ui/. Pangal inherits the core 15 native consonants of Proto-Meitei, expanding its phonetic inventory to include voiced segments due to historical Indo-Aryan and bilingual contact. The consonant speech sounds recorded in this dataset are- /p/, /pʰ/, /b/,/t/, /tʰ/, /d/, /k/, /kʰ/, /g/, /m/, /n/, /ŋ/, /r/, /s/, /h/, /l/, /j/, /t͡ʃ/, /d͡ʒ/, /w/. Voiced aspirated consonants are absent in native elements. The syllable structure is either CV (Consonant-Vowel) or CVC (Consonant-Vowel-Consonant), with no complex coda. Pangal utilises a binary tonal contrast (Level versus Falling tone) to distinguish lexical meanings. A unique phonological rule involves an extended voicing process: in standard Meitei, voiceless stops undergo voicing after voiced segments (e.g., the suffix -pa becomes -ba), whereas in the Pangal variety, this rule applies even when preceded by an unaspirated voiceless stop (e.g., standard kap-pa shifts toward a more heavily voiced, localised pronunciation such as kabba).

Morphologically, Pangal is highly agglutinative and predominantly suffixing. Words are formed by attaching grammatical suffixes to a stable root. For example, the noun ‘ima’ (mother) can take the plural suffix ‘-sing’ to become ‘ima-sing’ (mothers). Similarly, the verb root ‘tʃa’ (eat) can be combined with the past tense marker ‘-re’ to form ‘tʃa-re’ (ate), and with the incomplete negation suffix ‘-dari’ to form ‘tʃa-dari’ (has not eaten yet). There is no agreement marking; grammatical gender, number, and person do not trigger structural modifications on the verb. Gender is indicated lexically, using markers such as -nupa (male) and -nupi (female), as in ‘matʃa-nupa’ (son) and ‘matʃa-nupi’ (daughter). The language exhibits a simplified lexical category system, with only two major open word classes: nouns and verbs. Adjectives and adverbs are not independent categories but are derived morphologically from verbal roots. For instance, the adjective meaning ‘long’ can be formed from the verb root ‘saŋ’ (to be long) by adding prefix ə- and the participial suffix -ba resulting in ə-saŋ-ba (long).

The Pangal variety follows a Subject-Object-Verb (SOV) word order. For example: /əina (I- NOM) tʃak (rice) tʃari (eating)/ which means ‘I am eating food’. Despite this structural baseline, the language exhibits pragmatic flexibility. Important thematic elements may be fronted to the beginning of a sentence for emphasis. For instance, to highlight the object through fronting, a speaker might utilise an OSV structure such as /tʃak əina tʃari/ (Food, I am eating), thereby focusing primary attention on the object.

About Project BoLI

Project BoLI is the flagship project of UnReaL-TecE LLP, which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the project website.

Ethical Considerations, Consent, IPR and Attribution

Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.

Project BoLI - Data Governance Policy

BoLI Ethics Principle & Pledge

Project BoLI - Digital Consent Form

Project BoLI - Field Recording of Oral Consent

Project BoLI - TnC

Dataset Access

Full dataset for the language can be browsed and queried on our app. The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

full CC BY-NC-SA 4.0 Open resource → Explore relationships →

dataset

BoLI Khasi Translation 1

The Dataset

The current data preview of Khasi is being released as part of the Project BoLI. This preview is a reflection of the full dataset and consists of the following -

  1. Speech Recordings of 100 sentences in the language.
  2. Transcriptions in IPA and native script(s).
  3. Translations in English (which also act as prompts for the translation sentences).
  4. Detailed speaker metadata, including their demographic, educational and linguistic profile.
  5. Prompt in English and Hindi.

The full dataset contains the following -

  1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones.
  2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks.
  3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations.
  4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc.
  5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors.

Dataset Preparation

The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language.

The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the guidelines for the project using MATra Lab, a part of the LiFE Suite Ecosystem, developed by Unreal Tece LLP.

Project BoLI Guidelines

About Khasi

Khasi belongs to the Austroasiatic language family, specifically classified within the Khasic branch.

Speakers

According to the 2011 Census of India, there are approximately 1,431,344 speakers of Khasi in India.

Distribution in India

The language is primarily spoken in the state of Meghalaya, particularly within the East Khasi Hills, West Khasi Hills, South West Khasi Hills, Eastern West Khasi Hills, and Ri-Bhoi districts. Significant speaker communities are also found in the neighbouring state of Assam.

Major Grammatical Features

  • Phonology: Khasi is a non-tonal language. It features a moderately large inventory of consonants, including aspirated stops, and allows complex initial consonant clusters. It contrasts short and long vowels.
  • Morphology: Khasi is largely analytical, though it employs a rich set of derivational prefixes (such as the causative pyn-) and infixes. Nouns are obligatorily marked for grammatical gender using four proclitic articles: u (masculine), ka (feminine), i (diminutive/affectionate), and ki (plural).
  • Syntax: The basic word order is strictly Subject-Verb-Object (SVO). It is a head-initial language, meaning that prepositions are used instead of postpositions, and nouns typically precede their modifying adjectives and genitive phrases.

Further reading

About Project BoLI

Project BoLI is the flagship project of UnReaL-TecE LLP, which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the project website.

Ethical Considerations, Consent, IPR and Attribution

Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.

Project BoLI - Data Governance Policy

BoLI Ethics Principle & Pledge

Project BoLI - Digital Consent Form

Project BoLI - Field Recording of Oral Consent

Project BoLI - TnC

Dataset Access

Full dataset for the language can be browsed and queried on our app. The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

full CC BY-NC-SA 4.0 Open resource → Explore relationships →

dataset

BoLI Khasi Translation 2

The Dataset

The current data preview of Khasi is being released as part of the Project BoLI. This preview is a reflection of the full dataset and consists of the following -

  1. Speech Recordings of 100 sentences in the language.
  2. Transcriptions in IPA and native script(s).
  3. Translations in English (which also act as prompts for the translation sentences).
  4. Detailed speaker metadata, including their demographic, educational and linguistic profile.
  5. Prompt in English and Hindi.

The full dataset contains the following -

  1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones.
  2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks.
  3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations.
  4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc.
  5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors.

Dataset Preparation

The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language.

The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the guidelines for the project using MATra Lab, a part of the LiFE Suite Ecosystem, developed by Unreal Tece LLP.

Project BoLI Guidelines

About Khasi

Khasi belongs to the Austroasiatic language family, specifically classified within the Khasic branch.

Speakers

According to the 2011 Census of India, there are approximately 1,431,344 speakers of Khasi in India.

Distribution in India

The language is primarily spoken in the state of Meghalaya, particularly within the East Khasi Hills, West Khasi Hills, South West Khasi Hills, Eastern West Khasi Hills, and Ri-Bhoi districts. Significant speaker communities are also found in the neighbouring state of Assam.

Major Grammatical Features

  • Phonology: Khasi is a non-tonal language. It features a moderately large inventory of consonants, including aspirated stops, and allows complex initial consonant clusters. It contrasts short and long vowels.
  • Morphology: Khasi is largely analytical, though it employs a rich set of derivational prefixes (such as the causative pyn-) and infixes. Nouns are obligatorily marked for grammatical gender using four proclitic articles: u (masculine), ka (feminine), i (diminutive/affectionate), and ki (plural).
  • Syntax: The basic word order is strictly Subject-Verb-Object (SVO). It is a head-initial language, meaning that prepositions are used instead of postpositions, and nouns typically precede their modifying adjectives and genitive phrases.

Further reading

About Project BoLI

Project BoLI is the flagship project of UnReaL-TecE LLP, which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the project website.

Ethical Considerations, Consent, IPR and Attribution

Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.

Project BoLI - Data Governance Policy

BoLI Ethics Principle & Pledge

Project BoLI - Digital Consent Form

Project BoLI - Field Recording of Oral Consent

Project BoLI - TnC

Dataset Access

Full dataset for the language can be browsed and queried on our app. The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

full CC BY-NC-SA 4.0 Open resource → Explore relationships →

dataset

BoLI Khasi Translation Benchmark

The Dataset

The current data preview of Khasi is being released as part of the Project BoLI. This preview is a reflection of the full dataset and consists of the following -

  1. Speech Recordings of 50 sentences in the language.
  2. Transcriptions in IPA and native script(s).
  3. Translations in English (which also act as prompts for the translation sentences).
  4. Detailed speaker metadata, including their demographic, educational and linguistic profile.
  5. Prompt in English and Hindi.

The full dataset contains the following -

  1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones.
  2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks.
  3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations.
  4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc.
  5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors.

Dataset Preparation

The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language.

The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the guidelines for the project using MATra Lab, a part of the LiFE Suite Ecosystem, developed by Unreal Tece LLP.

Project BoLI Guidelines

About Khasi

Khasi belongs to the Austroasiatic language family, specifically classified within the Khasic branch.

Speakers

According to the 2011 Census of India, there are approximately 1,431,344 speakers of Khasi in India.

Distribution in India

The language is primarily spoken in the state of Meghalaya, particularly within the East Khasi Hills, West Khasi Hills, South West Khasi Hills, Eastern West Khasi Hills, and Ri-Bhoi districts. Significant speaker communities are also found in the neighbouring state of Assam.

Major Grammatical Features

  • Phonology: Khasi is a non-tonal language. It features a moderately large inventory of consonants, including aspirated stops, and allows complex initial consonant clusters. It contrasts short and long vowels.
  • Morphology: Khasi is largely analytical, though it employs a rich set of derivational prefixes (such as the causative pyn-) and infixes. Nouns are obligatorily marked for grammatical gender using four proclitic articles: u (masculine), ka (feminine), i (diminutive/affectionate), and ki (plural).
  • Syntax: The basic word order is strictly Subject-Verb-Object (SVO). It is a head-initial language, meaning that prepositions are used instead of postpositions, and nouns typically precede their modifying adjectives and genitive phrases.

Further reading

About Project BoLI

Project BoLI is the flagship project of UnReaL-TecE LLP, which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the project website.

Ethical Considerations, Consent, IPR and Attribution

Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.

Project BoLI - Data Governance Policy

BoLI Ethics Principle & Pledge

Project BoLI - Digital Consent Form

Project BoLI - Field Recording of Oral Consent

Project BoLI - TnC

Dataset Access

Full dataset for the language can be browsed and queried on our app. The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

full CC BY-NC-SA 4.0 Open resource → Explore relationships →

dataset

BoLI Garo Translation 2

The Dataset

The current data preview of Garo is being released as part of the Project BoLI. This preview is a reflection of the full dataset and consists of the following -

  1. Speech Recordings of 100 sentences in the language.
  2. Transcriptions in IPA and native script(s).
  3. Translations in English (which also act as prompts for the translation sentences).
  4. Detailed speaker metadata, including their demographic, educational and linguistic profile.
  5. Prompt in English and Hindi.

The full dataset contains the following -

  1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones.
  2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks.
  3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations.
  4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc.
  5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors.

### Dataset Preparation

The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language.

The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the guidelines for the project using MATra Lab, a part of the LiFE Suite Ecosystem, developed by Unreal Tece LLP.

Project BoLI Guidelines

### About Garo

Garo belongs to the Tibeto-Burman language family, situated within the Bodo-Garo branch of the Sal subfamily.

Speakers

According to the 2011 Census of India, there are approximately 1,145,323 native Garo speakers in India.

Distribution in India

The principal concentration of Garo speakers is in the state of Meghalaya, particularly within the East, West, North, South, and Southwest Garo Hills districts. Significant communities also reside in adjacent areas of Assam (including the Goalpara, Kamrup, and Karbi Anglong districts), Tripura, and West Bengal.

Major grammatical features

  • Phonology: Garo has a relatively simple vowel system (/a, e, i, o, u/) and a consonant inventory characterized by a distinction between aspirated and unaspirated voiceless stops. A defining feature is the glottal stop /ʔ/ (often orthographically represented as q or an apostrophe), which frequently appears in syllable-coda positions. Unlike many Tibeto-Burman languages, Garo is generally considered non-tonal.
  • Morphology: It is highly agglutinative, relying heavily on suffixation. Noun phrases employ a rich system of numeral classifiers (e.g., sak for humans, gong for flat objects) that must accompany numerals. Nouns are marked for case (nominative, accusative, genitive, dative, locative, and instrumental) via postpositional clitics. Verbs are morphologically complex, inflecting for tense, aspect, mood, and negation through chain-like suffixation.
  • Syntax: Garo exhibits a rigid Subject-Object-Verb (SOV) basic word order. It is a postpositional language where modifiers, such as adjectives, typically follow the noun they modify, while genitives and numeral classifiers precede the noun head.

Further reading

### About Project BoLI

Project BoLI is the flagship project of UnReaL-TecE LLP, which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the project website.

### Ethical Considerations, Consent, IPR and Attribution

Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.

Project BoLI - Data Governance Policy

BoLI Ethics Principle & Pledge

Project BoLI - Digital Consent Form

Project BoLI - Field Recording of Oral Consent

Project BoLI - TnC

### Dataset Access

Full dataset for the language can be browsed and queried on our app. The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

full CC BY-NC-SA 4.0 Open resource → Explore relationships →

dataset

BoLI Garo Translation 1

The Dataset

The current data preview of Garo is being released as part of the Project BoLI. This preview is a reflection of the full dataset and consists of the following -

  1. Speech Recordings of 100 sentences in the language.
  2. Transcriptions in IPA and native script(s).
  3. Translations in English (which also act as prompts for the translation sentences).
  4. Detailed speaker metadata, including their demographic, educational and linguistic profile.
  5. Prompt in English and Hindi.

The full dataset contains the following -

  1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones.
  2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks.
  3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations.
  4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc.
  5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors.

### Dataset Preparation

The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language.

The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the guidelines for the project using MATra Lab, a part of the LiFE Suite Ecosystem, developed by Unreal Tece LLP.

Project BoLI Guidelines

### About Garo

Garo belongs to the Tibeto-Burman language family, situated within the Bodo-Garo branch of the Sal subfamily.

Speakers

According to the 2011 Census of India, there are approximately 1,145,323 native Garo speakers in India.

Distribution in India

The principal concentration of Garo speakers is in the state of Meghalaya, particularly within the East, West, North, South, and Southwest Garo Hills districts. Significant communities also reside in adjacent areas of Assam (including the Goalpara, Kamrup, and Karbi Anglong districts), Tripura, and West Bengal.

Major grammatical features

  • Phonology: Garo has a relatively simple vowel system (/a, e, i, o, u/) and a consonant inventory characterized by a distinction between aspirated and unaspirated voiceless stops. A defining feature is the glottal stop /ʔ/ (often orthographically represented as q or an apostrophe), which frequently appears in syllable-coda positions. Unlike many Tibeto-Burman languages, Garo is generally considered non-tonal.
  • Morphology: It is highly agglutinative, relying heavily on suffixation. Noun phrases employ a rich system of numeral classifiers (e.g., sak for humans, gong for flat objects) that must accompany numerals. Nouns are marked for case (nominative, accusative, genitive, dative, locative, and instrumental) via postpositional clitics. Verbs are morphologically complex, inflecting for tense, aspect, mood, and negation through chain-like suffixation.
  • Syntax: Garo exhibits a rigid Subject-Object-Verb (SOV) basic word order. It is a postpositional language where modifiers, such as adjectives, typically follow the noun they modify, while genitives and numeral classifiers precede the noun head.

Further reading

### About Project BoLI

Project BoLI is the flagship project of UnReaL-TecE LLP, which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the project website.

### Ethical Considerations, Consent, IPR and Attribution

Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.

Project BoLI - Data Governance Policy

BoLI Ethics Principle & Pledge

Project BoLI - Digital Consent Form

Project BoLI - Field Recording of Oral Consent

Project BoLI - TnC

### Dataset Access

Full dataset for the language can be browsed and queried on our app. The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

full CC BY-NC-SA 4.0 Open resource → Explore relationships →

dataset

BoLI Tulu Translation

The Dataset

The current data preview of Tulu is being released as part of the Project BoLI. This preview is a reflection of the full dataset and consists of the following -

  1. Speech Recordings of 200 sentences in the language.
  2. Transcriptions in IPA and Bangali, English
  3. Translations in English (which also act as prompts for the translation sentences)
  4. Detailed speaker metadata, including their demographic, educational and linguistic profile.
  5. Prompt in English and Hindi.

The full dataset contains the following -

  1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones.
  2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks.
  3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations.
  4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc.
  5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors.

Dataset Preparation

The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language.

The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the guidelines for the project using MATra Lab, a part of the LiFE Suite Ecosystem, developed by Unreal Tece LLP.

Project BoLI Guidelines

About Tulu

Tulu belongs to the Dravidian language family, specifically classified within the South Dravidian branch.

Speakers

According to the 2011 Census of India, there are approximately 1,846,427 native Tulu speakers in India.

Distribution in India

Tulu is principally spoken in the southwestern coastal region of India, a belt traditionally referred to as Tulu Nadu. This region primarily encompasses the Dakshina Kannada and Udupi districts in the state of Karnataka, as well as the northern part of the Kasaragod district in Kerala.

Major grammatical features

  • Phonology: Tulu possesses a typical Dravidian phonemic inventory, including a series of retroflex consonants (/ʈ/, /ɖ/, /ɳ/, /ɭ/). It is highly notable for its vowel system, which includes the close back unrounded vowel (/ɯ/) in addition to the close back rounded vowel (/u/).
  • Morphology: Highly agglutinative. Nouns are marked for number (singular and plural) and decline for several cases, including nominative, accusative, dative, genitive, locative, and instrumental/sociative. Tulu distinguishes three grammatical genders in the singular (masculine, feminine, and neuter), though the masculine and feminine collapse into a common "rational" class in the plural. Verbs conjugate for tense (past, present, future), mood, and person-number-gender agreement.
  • Syntax: Tulu strictly adheres to a Subject-Object-Verb (SOV) default word order. It is a left-branching language, meaning modifiers and relative clauses typically precede their head nouns, and postpositions are used instead of prepositions.

About Project BoLI

Project BoLI is the flagship project of UnReaL-TecE LLP, which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the project website.

Ethical Considerations, Consent, IPR and Attribution

Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.

Project BoLI - Data Governance Policy

BoLI Ethics Principle & Pledge

Project BoLI - Digital Consent Form

Project BoLI - Field Recording of Oral Consent

Project BoLI - TnC

Dataset Access

Full dataset for the language can be browsed and queried here. The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

full CC BY-NC-SA 4.0 Open resource → Explore relationships →

dataset

BoLI Tulu Narration

The Dataset

The current data preview of Tulu is being released as part of the Project BoLI. This preview is a reflection of the full dataset and consists of the following -

  1. Speech Recordings of 10 narrations in the language.
  2. Transcriptions in IPA and Bangali, English
  3. Translations in English (which also act as prompts for the translation sentences)
  4. Detailed speaker metadata, including their demographic, educational and linguistic profile.
  5. Prompt in English and Hindi.

The full dataset contains the following -

  1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones.
  2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks.
  3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations.
  4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc.
  5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors.

Dataset Preparation

The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language.

The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the guidelines for the project using MATra Lab, a part of the LiFE Suite Ecosystem, developed by Unreal Tece LLP.

Project BoLI Guidelines

About Tulu

Tulu belongs to the Dravidian language family, specifically classified within the South Dravidian branch.

Speakers

According to the 2011 Census of India, there are approximately 1,846,427 native Tulu speakers in India.

Distribution in India

Tulu is principally spoken in the southwestern coastal region of India, a belt traditionally referred to as Tulu Nadu. This region primarily encompasses the Dakshina Kannada and Udupi districts in the state of Karnataka, as well as the northern part of the Kasaragod district in Kerala.

Major grammatical features

  • Phonology: Tulu possesses a typical Dravidian phonemic inventory, including a series of retroflex consonants (/ʈ/, /ɖ/, /ɳ/, /ɭ/). It is highly notable for its vowel system, which includes the close back unrounded vowel (/ɯ/) in addition to the close back rounded vowel (/u/).
  • Morphology: Highly agglutinative. Nouns are marked for number (singular and plural) and decline for several cases, including nominative, accusative, dative, genitive, locative, and instrumental/sociative. Tulu distinguishes three grammatical genders in the singular (masculine, feminine, and neuter), though the masculine and feminine collapse into a common "rational" class in the plural. Verbs conjugate for tense (past, present, future), mood, and person-number-gender agreement.
  • Syntax: Tulu strictly adheres to a Subject-Object-Verb (SOV) default word order. It is a left-branching language, meaning modifiers and relative clauses typically precede their head nouns, and postpositions are used instead of prepositions.

About Project BoLI

Project BoLI is the flagship project of UnReaL-TecE LLP, which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the project website.

Ethical Considerations, Consent, IPR and Attribution

Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.

Project BoLI - Data Governance Policy

BoLI Ethics Principle & Pledge

Project BoLI - Digital Consent Form

Project BoLI - Field Recording of Oral Consent

Project BoLI - TnC

Dataset Access

Full dataset for the language can be browsed and queried here. The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

full CC BY-NC-SA 4.0 Open resource → Explore relationships →

dataset

BoLI Toto Translation

The Dataset

The current data preview of Toto is being released as part of the Project BoLI. This preview is a reflection of the full dataset and consists of the following -

  1. Speech Recordings of 200 sentences in the language.
  2. Transcriptions in IPA and native script
  3. Translations in English (which also act as prompts for the translation sentences)
  4. Detailed speaker metadata, including their demographic, educational and linguistic profile.
  5. Prompt in English and Hindi.

The full dataset contains the following -

  1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones.
  2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks.
  3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations.
  4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc.
  5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors.

Dataset Preparation

The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language.

The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the guidelines for the project using MATra Lab, a part of the LiFE Suite Ecosystem, developed by Unreal Tece LLP.

Project BoLI Guidelines

About Toto

Toto (ISO 639-3: txo, Glottocode: toto1302) is a member of the Tibeto-Burman language family. Within this family, it is generally classified under the Sal group or grouped alongside Dhimal in the tentative Dhimalish branch, though its precise sub-classification remains subject to ongoing academic debate due to limited historical data.

According to the 2011 Census of India, there are approximately 1,600 members of Toto tribe who identify themseleves as Toto speakers. However, this total community-wide population is distinct from the small, select group of speakers which actually speak the language. The language is currently classified as critically endangered due to the small size of the community and the increasing influence of neighbouring major languages.

Distribution in India

Toto is spoken almost exclusively in Totopara, a small village enclave located in the Alipurduar district (formerly part of the Jalpaiguri district) of West Bengal, India, situated along the border with Bhutan.

Major Grammatical Features

  • Phonology: Toto features a relatively small inventory of vowels and consonants. It is widely considered to have a tonal or register distinction, though the exact phonological behaviour of its pitch system remains under-documented and uncertain.
  • Morphology: The language is predominantly agglutinating. Nouns are inflected for case using postpositions (including ergative/agentive, genitive, accusative, dative, and locative markers). Verbs inflect for tense, aspect, and negation, although the complexity of verbal agreement patterns appears less pronounced than in other Tibeto-Burman languages of the region.
  • Syntax: The basic word order in Toto is Subject-Object-Verb (SOV). It exhibits head-final typological features, with modifiers such as genitives and relative clauses typically preceding the head noun.

About Project BoLI

Project BoLI is the flagship project of UnReaL-TecE LLP, which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the project website.

Ethical Considerations, Consent, IPR and Attribution

Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.

Project BoLI - Data Governance Policy

BoLI Ethics Principle & Pledge

Project BoLI - Digital Consent Form

Project BoLI - Field Recording of Oral Consent

Project BoLI - TnC

Dataset Access

Full dataset for the language can be browsed and queried here. The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

full CC BY-NC-SA 4.0 Open resource → Explore relationships →

dataset

BoLI Toto Narration

The Dataset

The current data preview of Toto is being released as part of the Project BoLI. This preview is a reflection of the full dataset and consists of the following -

  1. Speech Recordings of 10 narrations in the language.
  2. Transcriptions in IPA and native script
  3. Translations in English (which also act as prompts for the translation sentences)
  4. Detailed speaker metadata, including their demographic, educational and linguistic profile.
  5. Prompt in English and Hindi.

The full dataset contains the following -

  1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones.
  2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks.
  3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations.
  4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc.
  5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors.

Dataset Preparation

The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language.

The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the guidelines for the project using MATra Lab, a part of the LiFE Suite Ecosystem, developed by Unreal Tece LLP.

Project BoLI Guidelines

About Toto

Toto (ISO 639-3: txo, Glottocode: toto1302) is a member of the Tibeto-Burman language family. Within this family, it is generally classified under the Sal group or grouped alongside Dhimal in the tentative Dhimalish branch, though its precise sub-classification remains subject to ongoing academic debate due to limited historical data.

According to the 2011 Census of India, there are approximately 1,600 members of Toto tribe who identify themseleves as Toto speakers. However, this total community-wide population is distinct from the small, select group of speakers which actually speak the language. The language is currently classified as critically endangered due to the small size of the community and the increasing influence of neighbouring major languages.

Distribution in India

Toto is spoken almost exclusively in Totopara, a small village enclave located in the Alipurduar district (formerly part of the Jalpaiguri district) of West Bengal, India, situated along the border with Bhutan.

Major Grammatical Features

  • Phonology: Toto features a relatively small inventory of vowels and consonants. It is widely considered to have a tonal or register distinction, though the exact phonological behaviour of its pitch system remains under-documented and uncertain.
  • Morphology: The language is predominantly agglutinating. Nouns are inflected for case using postpositions (including ergative/agentive, genitive, accusative, dative, and locative markers). Verbs inflect for tense, aspect, and negation, although the complexity of verbal agreement patterns appears less pronounced than in other Tibeto-Burman languages of the region.
  • Syntax: The basic word order in Toto is Subject-Object-Verb (SOV). It exhibits head-final typological features, with modifiers such as genitives and relative clauses typically preceding the head noun.

About Project BoLI

Project BoLI is the flagship project of UnReaL-TecE LLP, which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the project website.

Ethical Considerations, Consent, IPR and Attribution

Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.

Project BoLI - Data Governance Policy

BoLI Ethics Principle & Pledge

Project BoLI - Digital Consent Form

Project BoLI - Field Recording of Oral Consent

Project BoLI - TnC

Dataset Access

Full dataset for the language can be browsed and queried here. The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

full CC BY-NC-SA 4.0 Open resource → Explore relationships →

dataset

BoLI Erava Narration

The Dataset

The current data preview of Erava is being released as part of the Project BoLI. This preview is a reflection of the full dataset and consists of the following -

  1. Speech Recordings of 10 narrations in the language.
  2. Transcriptions in IPA and native script(s).
  3. Translations in English (which also act as prompts for the translation sentences).
  4. Detailed speaker metadata, including their demographic, educational and linguistic profile.
  5. Prompt in English and Hindi.

The full dataset contains the following -

  1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones.
  2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks.
  3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations.
  4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc.
  5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors.

Dataset Preparation

The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language.

The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the guidelines for the project using MATra Lab, a part of the LiFE Suite Ecosystem, developed by Unreal Tece LLP.

Project BoLI Guidelines

About Erava

Erava (often identified with or closely related to Ravula) belongs to the Dravidian language family, specifically classified within the Southern Dravidian subgroup.

Speakers

The total number of Erava speakers remains highly uncertain due to varying classifications under the "Yerava" or "Ravula" labels. However, scholarly estimates and tribal surveys place the ethnic speaker population at approximately 25,000 to 30,000, as indexed in reports like the Census of India 2011.

Distribution in India

Erava is primarily spoken in the Southern region of India. Its principal concentration is in the Kodagu (Coorg) district of Karnataka, with some speakers also residing in the neighboring districts of Kerala.

Major grammatical features

  • Phonology: Erava features a typical South Dravidian phonological system, including a distinction between short and long vowels, and a consonant inventory that features retroflex stops (/ʈ/, /ɖ/), nasals (/ɳ/), and laterals (/ɭ/).
  • Morphology: It is a highly agglutinative language. Nouns inflect for case (including nominative, accusative, dative, genitive, and locative) and number. Verbs inflect for tense (distinguishing past and non-past), mood, and aspect, typically agreeing with the subject in person, number, and gender.
  • Syntax: The basic word order is Subject-Object-Verb (SOV). It is a head-final language employing postpositions rather than prepositions, and relative clauses precede the nouns they modify.

About Project BoLI

Project BoLI is the flagship project of UnReaL-TecE LLP, which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the project website.

Ethical Considerations, Consent, IPR and Attribution

Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.

Project BoLI - Data Governance Policy

BoLI Ethics Principle & Pledge

Project BoLI - Digital Consent Form

Project BoLI - Field Recording of Oral Consent

Project BoLI - TnC

Dataset Access

Full dataset for the language can be browsed and queried here. The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

full CC BY-NC-SA 4.0 Open resource → Explore relationships →

dataset

BoLI Ho Translation

The Dataset

The current data preview of Ho is being released as part of the Project BoLI. This preview is a reflection of the full dataset and consists of the following -

  1. Speech Recordings of 200 sentences in the language.
  2. Transcriptions in IPA and native script(s).
  3. Translations in English (which also act as prompts for the translation sentences).
  4. Detailed speaker metadata, including their demographic, educational and linguistic profile.
  5. Prompt in English and Hindi.

The full dataset contains the following -

  1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones.
  2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks.
  3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations.
  4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc.
  5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors.

Dataset Preparation

The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language.

The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the guidelines for the project using MATra Lab, a part of the LiFE Suite Ecosystem, developed by Unreal Tece LLP.

Project BoLI Guidelines

About Ho

Ho belongs to the Austro-Asiatic language family, specifically situated within the North Munda branch of the Munda sub-family.

Speakers

According to the 2011 Census of India, there are approximately 1.4 million speakers of the Ho language.

Distribution in India

The language is primarily spoken in the state of Jharkhand (especially concentrated in the West Singhbhum and East Singhbhum districts). It is also widely spoken in the adjoining districts of northern Odisha (such as Mayurbhanj, Keonjhar, and Sundargarh) and parts of West Bengal.

Major grammatical features

  • Phonology: The phonemic inventory of Ho includes a distinction between aspirated and unaspirated stops, retroflex consonants, and characteristic Austro-Asiatic glottalized or checked stops (often occurring syllable-finally). Vowel nasalization and length can also carry contrastive functional weight.
  • Morphology: Ho is a highly agglutinative language. Nouns are marked for three grammatical numbers (singular, dual, and plural) and inflected for grammatical case. Verbs are highly complex, incorporating pronominal markers for both the subject and object, alongside suffixes indicating tense, aspect, and mood.
  • Syntax: The default word order is Subject-Object-Verb (SOV). Ho is head-final, utilizing postpositions rather than prepositions, with modifiers generally preceding the nouns they qualify.

About Project BoLI

Project BoLI is the flagship project of UnReaL-TecE LLP, which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the project website.

Ethical Considerations, Consent, IPR and Attribution

Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.

Project BoLI - Data Governance Policy

BoLI Ethics Principle & Pledge

Project BoLI - Digital Consent Form

Project BoLI - Field Recording of Oral Consent

Project BoLI - TnC

Dataset Access

Full dataset for the language can be browsed and queried on our app. The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

full CC BY-NC-SA 4.0 Open resource → Explore relationships →