Dataset created under the Speech Datasets and Models for Tibeto-Burman Languages (SpeeD-TB), sponsored under Mission Bhashini by Ministry of Electronics and Information Technology (MEITY), Govt of India. The project aimed to create 1,200 hours of speech dataset and ASR models for 6 underresourced, tribal Tibeto-Burman languages of India speoken in Eastern and North-Eastern parts of India.
Project BoLI, the flagship initiative of UnReaL-TecE LLP, is building high-quality datasets to train and benchmark next-generation AI tasks—from speech-to-text to advanced grammatical reasoning—across more than 1,000 underserved Indian languages and dialects. We are pioneering a globally unique data governance framework where all released datasets, derivative assets, and downstream models are permanently co-owned by the contributing communities. Under our groundbreaking BoLI License, commercial data usage is strictly treated as a revocable license rather than a transfer of ownership, ensuring total accountability in the era of LLMs.
Project BoLI, the flagship initiative of UnReaL-TecE LLP, is building high-quality datasets to train and benchmark next-generation AI tasks—from speech-to-text to advanced grammatical reasoning—across more than 1,000 underserved Indian languages and dialects. We are pioneering a globally unique data governance framework where all released datasets, derivative assets, and downstream models are permanently co-owned by the contributing communities. Under our groundbreaking BoLI License, commercial data usage is strictly treated as a revocable license rather than a transfer of ownership, ensuring total accountability in the era of LLMs.
Dataset created under the Speech Datasets and Models for Tibeto-Burman Languages (SpeeD-TB), sponsored under Mission Bhashini by Ministry of Electronics and Information Technology (MEITY), Govt of India. The project aimed to create 1,200 hours of speech dataset and ASR models for 6 underresourced, tribal Tibeto-Burman languages of India speoken in Eastern and North-Eastern parts of India.
Dataset created under the Speech Datasets and Models for Tibeto-Burman Languages (SpeeD-TB), sponsored under Mission Bhashini by Ministry of Electronics and Information Technology (MEITY), Govt of India. The project aimed to create 1,200 hours of speech dataset and ASR models for 6 underresourced, tribal Tibeto-Burman languages of India speoken in Eastern and North-Eastern parts of India.
## About Toto
Toto is an under-resourced and critically endangered Tibeto-Burman language spoken by a small community in Totopara village, Alipurduar district, West Bengal, India,
with ** fewer than 1,000 speakers** (and a significantly smaller number of people proficient in the language). The language belongs to the Dhimalish group of languages and is
closely related to Dhimal, another language spoken in Northern West Bengal.
## Dataset Description
The Toto Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website")
funded under Mission Bhashini, is a transcribed speech corpus of Toto.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Bengali script**, making it
the **first and largest resource for the language** that not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the Toto language by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
The rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby giving extensive coverage.
We have also used a variety of elicitation methods for collecting the data, including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to rich metadata, including demographic and linguistic metadata of the
speakers, domains, elicitation methods and individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains the English-Toto parallel sentence dataset collected from the field.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment not only to fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.
## About Toto
Toto is an under-resourced and critically endangered Tibeto-Burman language spoken by a small community in Totopara village, Alipurduar district, West Bengal, India,
with ** fewer than 1,000 speakers** (and a significantly smaller number of people proficient in the language). The language belongs to the Dhimalish group of languages and is
closely related to Dhimal, another language spoken in Northern West Bengal.
## Dataset Description
The Toto Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website")
funded under Mission Bhashini, is a transcribed speech corpus of Toto.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Bengali script**, making it
the **first and largest resource for the language** that not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the Toto language by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
The rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby giving extensive coverage.
We have also used a variety of elicitation methods for collecting the data, including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to rich metadata, including demographic and linguistic metadata of the
speakers, domains, elicitation methods and individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains the Toto data collected from YouTube.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment not only to fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.
## About Toto
Toto is an under-resourced and critically endangered Tibeto-Burman language spoken by a small community in Totopara village, Alipurduar district, West Bengal, India,
with ** fewer than 1,000 speakers** (and a significantly smaller number of people proficient in the language). The language belongs to the Dhimalish group of languages and is
closely related to Dhimal, another language spoken in Northern West Bengal.
## Dataset Description
The Toto Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website")
funded under Mission Bhashini, is a transcribed speech corpus of Toto.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Bengali script**, making it
the **first and largest resource for the language** that not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the Toto language by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
The rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby giving extensive coverage.
We have also used a variety of elicitation methods for collecting the data, including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to rich metadata, including demographic and linguistic metadata of the
speakers, domains, elicitation methods and individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains the long-form lectures in Toto. Each lecture is approximately 30 minutes.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment not only to fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.
### About Chokri
Chokri (also known as Chakru, ISO 639-3: `nri`, Glottocode: `chok1243`) belongs to the Tibeto-Burman language family, specifically classified within the Angami-Pochuri branch of the Southern Tibeto-Burman subgroup.
### Speakers
According to the [2011 Census of India](https://censusindia.gov.in/census.website/data/census-tables), there are approximately 91,257 speakers of Chakru/Chokri in India.
### Distribution in India
Chokri is predominantly spoken in the northeastern state of Nagaland, India. It is concentrated heavily in the Phek district, particularly across the Pfütsero, Chetheba, and Chazouba administrative circles.
### Major grammatical features
* **Phonology:** Chokri is a tonal language characterised by a complex pitch/tone system that distinguishes lexical meaning. Its consonant inventory is notable for contrasting aspirated and unaspirated stops, alongside a series of voiceless sonorants (including voiceless nasals and laterals) typical of Angami-Pochuri languages.
* **Morphology:** The language is predominantly agglutinative. Morphological processes rely heavily on suffixation and prefixation for word formation. Nouns take possessive prefixes corresponding to person and number, and verbs are modified by a range of aspectual, modal, and directional markers.
* **Syntax:** Chokri exhibits a basic Subject-Object-Verb (SOV) constituent word order. It is a postpositional language where modifiers like numerals and demonstratives generally follow the head noun, while relative clauses can precede or follow the noun they modify.
## Dataset Description
The Chokri Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website")
funded under Mission Bhashini, is a transcribed speech corpus of Chokri.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Roman script**, making it
the **first and largest resource for the language** that not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the Toto language by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
Rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby, giving a large coverage.
We have also used a variety of elicitation methods for collecting the data including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to a rich metadats including demographic and linguistic metadata of the
speakers, domains, elicitation methods and to individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby, ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains English-Chokri translation data collected from the field using the translation method.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.
### About Chokri
Chokri (also known as Chakru, ISO 639-3: `nri`, Glottocode: `chok1243`) belongs to the Tibeto-Burman language family, specifically classified within the Angami-Pochuri branch of the Southern Tibeto-Burman subgroup.
### Speakers
According to the [2011 Census of India](https://censusindia.gov.in/census.website/data/census-tables), there are approximately 91,257 speakers of Chakru/Chokri in India.
### Distribution in India
Chokri is predominantly spoken in the northeastern state of Nagaland, India. It is concentrated heavily in the Phek district, particularly across the Pfütsero, Chetheba, and Chazouba administrative circles.
### Major grammatical features
* **Phonology:** Chokri is a tonal language characterised by a complex pitch/tone system that distinguishes lexical meaning. Its consonant inventory is notable for contrasting aspirated and unaspirated stops, alongside a series of voiceless sonorants (including voiceless nasals and laterals) typical of Angami-Pochuri languages.
* **Morphology:** The language is predominantly agglutinative. Morphological processes rely heavily on suffixation and prefixation for word formation. Nouns take possessive prefixes corresponding to person and number, and verbs are modified by a range of aspectual, modal, and directional markers.
* **Syntax:** Chokri exhibits a basic Subject-Object-Verb (SOV) constituent word order. It is a postpositional language where modifiers like numerals and demonstratives generally follow the head noun, while relative clauses can precede or follow the noun they modify.
## Dataset Description
The Chokri Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website")
funded under Mission Bhashini, is a transcribed speech corpus of Chokri.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Roman script**, making it
the **first and largest resource for the language** that not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the Toto language by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
Rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby, giving a large coverage.
We have also used a variety of elicitation methods for collecting the data including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to a rich metadats including demographic and linguistic metadata of the
speakers, domains, elicitation methods and to individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby, ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains English-Chokri translation data collected from the field using the translation method.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.
### About Chokri
Chokri (also known as Chakru, ISO 639-3: `nri`, Glottocode: `chok1243`) belongs to the Tibeto-Burman language family, specifically classified within the Angami-Pochuri branch of the Southern Tibeto-Burman subgroup.
### Speakers
According to the [2011 Census of India](https://censusindia.gov.in/census.website/data/census-tables), there are approximately 91,257 speakers of Chakru/Chokri in India.
### Distribution in India
Chokri is predominantly spoken in the northeastern state of Nagaland, India. It is concentrated heavily in the Phek district, particularly across the Pfütsero, Chetheba, and Chazouba administrative circles.
### Major grammatical features
* **Phonology:** Chokri is a tonal language characterised by a complex pitch/tone system that distinguishes lexical meaning. Its consonant inventory is notable for contrasting aspirated and unaspirated stops, alongside a series of voiceless sonorants (including voiceless nasals and laterals) typical of Angami-Pochuri languages.
* **Morphology:** The language is predominantly agglutinative. Morphological processes rely heavily on suffixation and prefixation for word formation. Nouns take possessive prefixes corresponding to person and number, and verbs are modified by a range of aspectual, modal, and directional markers.
* **Syntax:** Chokri exhibits a basic Subject-Object-Verb (SOV) constituent word order. It is a postpositional language where modifiers like numerals and demonstratives generally follow the head noun, while relative clauses can precede or follow the noun they modify.
## Dataset Description
The Chokri Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website")
funded under Mission Bhashini, is a transcribed speech corpus of Chokri.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Roman script**, making it
the **first and largest resource for the language** that not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the Toto language by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
Rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby, giving a large coverage.
We have also used a variety of elicitation methods for collecting the data including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to a rich metadats including demographic and linguistic metadata of the
speakers, domains, elicitation methods and to individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby, ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains English-Chokri translation data collected from the field using the translation method.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.
## About Bodo
Bodo belongs to the Tibeto-Burman language family, specifically belonging to the Bodo-Garo subgroup of the Sal language group.
According to the official [2011 Census of India](https://censusindia.gov.in/), there are approximately 1.48 million native Bodo speakers.
The principal concentration of Bodo speakers is in the state of Assam, particularly within the autonomous Bodoland Territorial Region. Smaller speech communities are also distributed across adjacent districts in the states of West Bengal, Meghalaya, and Nagaland.
## Dataset Description
The Bodo Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website"),
funded under Mission Bhashini, is a transcribed speech corpus of the language.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Roman script**, making it
** one of the largest speech resources for the language**, which not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the language and culture by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
The rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby giving extensive coverage.
We have also used a variety of elicitation methods for collecting the data, including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to rich metadata, including demographic and linguistic metadata of the
speakers, domains, elicitation methods and individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains an English-Bodo parallel dataset collected using the translation method.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.
## About Bodo
Bodo belongs to the Tibeto-Burman language family, specifically belonging to the Bodo-Garo subgroup of the Sal language group.
According to the official [2011 Census of India](https://censusindia.gov.in/), there are approximately 1.48 million native Bodo speakers.
The principal concentration of Bodo speakers is in the state of Assam, particularly within the autonomous Bodoland Territorial Region. Smaller speech communities are also distributed across adjacent districts in the states of West Bengal, Meghalaya, and Nagaland.
## Dataset Description
The Bodo Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website"),
funded under Mission Bhashini, is a transcribed speech corpus of the language.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Roman script**, making it
** one of the largest speech resources for the language**, which not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the language and culture by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
The rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby giving extensive coverage.
We have also used a variety of elicitation methods for collecting the data, including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to rich metadata, including demographic and linguistic metadata of the
speakers, domains, elicitation methods and individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains an English-Bodo parallel dataset collected using the translation method.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.
## About Bodo
Bodo belongs to the Tibeto-Burman language family, specifically belonging to the Bodo-Garo subgroup of the Sal language group.
According to the official [2011 Census of India](https://censusindia.gov.in/), there are approximately 1.48 million native Bodo speakers.
The principal concentration of Bodo speakers is in the state of Assam, particularly within the autonomous Bodoland Territorial Region. Smaller speech communities are also distributed across adjacent districts in the states of West Bengal, Meghalaya, and Nagaland.
## Dataset Description
The Bodo Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website"),
funded under Mission Bhashini, is a transcribed speech corpus of the language.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Roman script**, making it
** one of the largest speech resources for the language**, which not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the language and culture by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
The rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby giving extensive coverage.
We have also used a variety of elicitation methods for collecting the data, including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to rich metadata, including demographic and linguistic metadata of the
speakers, domains, elicitation methods and individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains an English-Bodo parallel dataset collected using the translation method.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.
## About Bodo
Bodo belongs to the Tibeto-Burman language family, specifically belonging to the Bodo-Garo subgroup of the Sal language group.
According to the official [2011 Census of India](https://censusindia.gov.in/), there are approximately 1.48 million native Bodo speakers.
The principal concentration of Bodo speakers is in the state of Assam, particularly within the autonomous Bodoland Territorial Region. Smaller speech communities are also distributed across adjacent districts in the states of West Bengal, Meghalaya, and Nagaland.
## Dataset Description
The Bodo Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website"),
funded under Mission Bhashini, is a transcribed speech corpus of the language.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Roman script**, making it
** one of the largest speech resources for the language**, which not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the language and culture by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
The rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby giving extensive coverage.
We have also used a variety of elicitation methods for collecting the data, including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to rich metadata, including demographic and linguistic metadata of the
speakers, domains, elicitation methods and individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains the full dataset for the
data collected from YouTube.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.
## About Toto
Toto is an under-resourced and critically endangered Tibeto-Burman language spoken by a small community in Totopara village, Alipurduar district, West Bengal, India,
with ** fewer than 1,000 speakers** (and a significantly smaller number of people proficient in the language). The language belongs to the Dhimalish group of languages and is
closely related to Dhimal, another language spoken in Northern West Bengal.
## Dataset Description
The Toto Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website")
funded under Mission Bhashini, is a transcribed speech corpus of Toto.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Bengali script**, making it
the **first and largest resource for the language** that not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the Toto language by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
The rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby giving extensive coverage.
We have also used a variety of elicitation methods for collecting the data, including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to rich metadata, including demographic and linguistic metadata of the
speakers, domains, elicitation methods and individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains the English-Toto parallel sentence dataset collected from the field.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment not only to fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.
## About Toto
Toto is an under-resourced and critically endangered Tibeto-Burman language spoken by a small community in Totopara village, Alipurduar district, West Bengal, India,
with ** fewer than 1,000 speakers** (and a significantly smaller number of people proficient in the language). The language belongs to the Dhimalish group of languages and is
closely related to Dhimal, another language spoken in Northern West Bengal.
## Dataset Description
The Toto Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website")
funded under Mission Bhashini, is a transcribed speech corpus of Toto.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Bengali script**, making it
the **first and largest resource for the language** that not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the Toto language by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
The rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby giving extensive coverage.
We have also used a variety of elicitation methods for collecting the data, including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to rich metadata, including demographic and linguistic metadata of the
speakers, domains, elicitation methods and individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains the Toto data collected from YouTube.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment not only to fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.
## About Toto
Toto is an under-resourced and critically endangered Tibeto-Burman language spoken by a small community in Totopara village, Alipurduar district, West Bengal, India,
with ** fewer than 1,000 speakers** (and a significantly smaller number of people proficient in the language). The language belongs to the Dhimalish group of languages and is
closely related to Dhimal, another language spoken in Northern West Bengal.
## Dataset Description
The Toto Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website")
funded under Mission Bhashini, is a transcribed speech corpus of Toto.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Bengali script**, making it
the **first and largest resource for the language** that not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the Toto language by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
The rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby giving extensive coverage.
We have also used a variety of elicitation methods for collecting the data, including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to rich metadata, including demographic and linguistic metadata of the
speakers, domains, elicitation methods and individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains the long-form lectures in Toto. Each lecture is approximately 30 minutes.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment not only to fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.
### About Chokri
Chokri (also known as Chakru, ISO 639-3: `nri`, Glottocode: `chok1243`) belongs to the Tibeto-Burman language family, specifically classified within the Angami-Pochuri branch of the Southern Tibeto-Burman subgroup.
### Speakers
According to the [2011 Census of India](https://censusindia.gov.in/census.website/data/census-tables), there are approximately 91,257 speakers of Chakru/Chokri in India.
### Distribution in India
Chokri is predominantly spoken in the northeastern state of Nagaland, India. It is concentrated heavily in the Phek district, particularly across the Pfütsero, Chetheba, and Chazouba administrative circles.
### Major grammatical features
* **Phonology:** Chokri is a tonal language characterised by a complex pitch/tone system that distinguishes lexical meaning. Its consonant inventory is notable for contrasting aspirated and unaspirated stops, alongside a series of voiceless sonorants (including voiceless nasals and laterals) typical of Angami-Pochuri languages.
* **Morphology:** The language is predominantly agglutinative. Morphological processes rely heavily on suffixation and prefixation for word formation. Nouns take possessive prefixes corresponding to person and number, and verbs are modified by a range of aspectual, modal, and directional markers.
* **Syntax:** Chokri exhibits a basic Subject-Object-Verb (SOV) constituent word order. It is a postpositional language where modifiers like numerals and demonstratives generally follow the head noun, while relative clauses can precede or follow the noun they modify.
## Dataset Description
The Chokri Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website")
funded under Mission Bhashini, is a transcribed speech corpus of Chokri.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Roman script**, making it
the **first and largest resource for the language** that not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the Toto language by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
Rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby, giving a large coverage.
We have also used a variety of elicitation methods for collecting the data including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to a rich metadats including demographic and linguistic metadata of the
speakers, domains, elicitation methods and to individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby, ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains English-Chokri translation data collected from the field using the translation method.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.
### About Chokri
Chokri (also known as Chakru, ISO 639-3: `nri`, Glottocode: `chok1243`) belongs to the Tibeto-Burman language family, specifically classified within the Angami-Pochuri branch of the Southern Tibeto-Burman subgroup.
### Speakers
According to the [2011 Census of India](https://censusindia.gov.in/census.website/data/census-tables), there are approximately 91,257 speakers of Chakru/Chokri in India.
### Distribution in India
Chokri is predominantly spoken in the northeastern state of Nagaland, India. It is concentrated heavily in the Phek district, particularly across the Pfütsero, Chetheba, and Chazouba administrative circles.
### Major grammatical features
* **Phonology:** Chokri is a tonal language characterised by a complex pitch/tone system that distinguishes lexical meaning. Its consonant inventory is notable for contrasting aspirated and unaspirated stops, alongside a series of voiceless sonorants (including voiceless nasals and laterals) typical of Angami-Pochuri languages.
* **Morphology:** The language is predominantly agglutinative. Morphological processes rely heavily on suffixation and prefixation for word formation. Nouns take possessive prefixes corresponding to person and number, and verbs are modified by a range of aspectual, modal, and directional markers.
* **Syntax:** Chokri exhibits a basic Subject-Object-Verb (SOV) constituent word order. It is a postpositional language where modifiers like numerals and demonstratives generally follow the head noun, while relative clauses can precede or follow the noun they modify.
## Dataset Description
The Chokri Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website")
funded under Mission Bhashini, is a transcribed speech corpus of Chokri.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Roman script**, making it
the **first and largest resource for the language** that not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the Toto language by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
Rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby, giving a large coverage.
We have also used a variety of elicitation methods for collecting the data including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to a rich metadats including demographic and linguistic metadata of the
speakers, domains, elicitation methods and to individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby, ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains English-Chokri translation data collected from the field using the translation method.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.
### About Chokri
Chokri (also known as Chakru, ISO 639-3: `nri`, Glottocode: `chok1243`) belongs to the Tibeto-Burman language family, specifically classified within the Angami-Pochuri branch of the Southern Tibeto-Burman subgroup.
### Speakers
According to the [2011 Census of India](https://censusindia.gov.in/census.website/data/census-tables), there are approximately 91,257 speakers of Chakru/Chokri in India.
### Distribution in India
Chokri is predominantly spoken in the northeastern state of Nagaland, India. It is concentrated heavily in the Phek district, particularly across the Pfütsero, Chetheba, and Chazouba administrative circles.
### Major grammatical features
* **Phonology:** Chokri is a tonal language characterised by a complex pitch/tone system that distinguishes lexical meaning. Its consonant inventory is notable for contrasting aspirated and unaspirated stops, alongside a series of voiceless sonorants (including voiceless nasals and laterals) typical of Angami-Pochuri languages.
* **Morphology:** The language is predominantly agglutinative. Morphological processes rely heavily on suffixation and prefixation for word formation. Nouns take possessive prefixes corresponding to person and number, and verbs are modified by a range of aspectual, modal, and directional markers.
* **Syntax:** Chokri exhibits a basic Subject-Object-Verb (SOV) constituent word order. It is a postpositional language where modifiers like numerals and demonstratives generally follow the head noun, while relative clauses can precede or follow the noun they modify.
## Dataset Description
The Chokri Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website")
funded under Mission Bhashini, is a transcribed speech corpus of Chokri.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Roman script**, making it
the **first and largest resource for the language** that not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the Toto language by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
Rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby, giving a large coverage.
We have also used a variety of elicitation methods for collecting the data including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to a rich metadats including demographic and linguistic metadata of the
speakers, domains, elicitation methods and to individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby, ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains English-Chokri translation data collected from the field using the translation method.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.
## About Bodo
Bodo belongs to the Tibeto-Burman language family, specifically belonging to the Bodo-Garo subgroup of the Sal language group.
According to the official [2011 Census of India](https://censusindia.gov.in/), there are approximately 1.48 million native Bodo speakers.
The principal concentration of Bodo speakers is in the state of Assam, particularly within the autonomous Bodoland Territorial Region. Smaller speech communities are also distributed across adjacent districts in the states of West Bengal, Meghalaya, and Nagaland.
## Dataset Description
The Bodo Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website"),
funded under Mission Bhashini, is a transcribed speech corpus of the language.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Roman script**, making it
** one of the largest speech resources for the language**, which not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the language and culture by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
The rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby giving extensive coverage.
We have also used a variety of elicitation methods for collecting the data, including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to rich metadata, including demographic and linguistic metadata of the
speakers, domains, elicitation methods and individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains an English-Bodo parallel dataset collected using the translation method.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.
## About Bodo
Bodo belongs to the Tibeto-Burman language family, specifically belonging to the Bodo-Garo subgroup of the Sal language group.
According to the official [2011 Census of India](https://censusindia.gov.in/), there are approximately 1.48 million native Bodo speakers.
The principal concentration of Bodo speakers is in the state of Assam, particularly within the autonomous Bodoland Territorial Region. Smaller speech communities are also distributed across adjacent districts in the states of West Bengal, Meghalaya, and Nagaland.
## Dataset Description
The Bodo Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website"),
funded under Mission Bhashini, is a transcribed speech corpus of the language.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Roman script**, making it
** one of the largest speech resources for the language**, which not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the language and culture by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
The rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby giving extensive coverage.
We have also used a variety of elicitation methods for collecting the data, including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to rich metadata, including demographic and linguistic metadata of the
speakers, domains, elicitation methods and individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains an English-Bodo parallel dataset collected using the translation method.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.
## About Bodo
Bodo belongs to the Tibeto-Burman language family, specifically belonging to the Bodo-Garo subgroup of the Sal language group.
According to the official [2011 Census of India](https://censusindia.gov.in/), there are approximately 1.48 million native Bodo speakers.
The principal concentration of Bodo speakers is in the state of Assam, particularly within the autonomous Bodoland Territorial Region. Smaller speech communities are also distributed across adjacent districts in the states of West Bengal, Meghalaya, and Nagaland.
## Dataset Description
The Bodo Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website"),
funded under Mission Bhashini, is a transcribed speech corpus of the language.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Roman script**, making it
** one of the largest speech resources for the language**, which not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the language and culture by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
The rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby giving extensive coverage.
We have also used a variety of elicitation methods for collecting the data, including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to rich metadata, including demographic and linguistic metadata of the
speakers, domains, elicitation methods and individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains an English-Bodo parallel dataset collected using the translation method.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.
## About Bodo
Bodo belongs to the Tibeto-Burman language family, specifically belonging to the Bodo-Garo subgroup of the Sal language group.
According to the official [2011 Census of India](https://censusindia.gov.in/), there are approximately 1.48 million native Bodo speakers.
The principal concentration of Bodo speakers is in the state of Assam, particularly within the autonomous Bodoland Territorial Region. Smaller speech communities are also distributed across adjacent districts in the states of West Bengal, Meghalaya, and Nagaland.
## Dataset Description
The Bodo Speech Dataset, developed as part of the
[Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB)](https://sites.google.com/view/speed-tb/ "SpeeD-TB Project Website"),
funded under Mission Bhashini, is a transcribed speech corpus of the language.
The full dataset comprises over **200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Roman script**, making it
** one of the largest speech resources for the language**, which not only enables building and evaluating voice models in
low-resource linguistic settings but also enables large-scale linguistic description and documentation of the language.
The audio data captures a diverse range of speakers across different age groups, genders, and education levels, ensuring variability in pronunciation,
speech patterns, and tone. It includes both spontaneous and read speech collected in naturalistic and semi-controlled environments, thereby reflecting
real-world linguistic usage. The transcriptions are carefully prepared and normalised to maintain consistency, supporting robust model training. This dataset
also contributes to the preservation and digital documentation of the language and culture by transforming oral knowledge into structured, machine-readable formats.
Almost 60% of the data in the corpus is included from domains of agriculture, education and science & technology.
The rest of the data is from varied domains including culture, lifecycle, sports, entertainment, healthcare and oral history, thereby giving extensive coverage.
We have also used a variety of elicitation methods for collecting the data, including translations, narrations, lectures, role-play, spontaneous conversations,
interviews and picture and video descriptions. The released dataset is meticulously mapped to rich metadata, including demographic and linguistic metadata of the
speakers, domains, elicitation methods and individual prompts. The audio included in the current dataset is already sliced at sentence level, thereby ready to be
integrated into the model training pipeline out-of-the-box.
The overall dataset of the project is collected over multiple phases and using multiple questionnaires. This repository contains the full dataset for the
data collected from YouTube.
## Ethical Considerations, IPR and Attribution
This repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset.
We believe this is the first step to move away from extractive data collection and use practices and ensure fairness in our treatment of the community members.
As such, we have listed all speakers and transcribers as Contributors to the dataset.
This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.