Dataset Showcase

BoLI Lambadi Translation 2

Whole project public CC BY-NC-SA 4.0 Low-resource LanguageUnderresourced LanguageUnreal TeceBoLIIndo-AryanLambadi

Research network

Explore this project in context

Follow its languages, people, related projects, and public data collections.

Select a connected item to inspect it and continue navigating.

About this project

The Dataset

The current data preview of Lambadi is being released as part of the Project BoLI. This preview is a reflection of the full dataset and consists of the following -

  1. Speech Recordings of 100 sentences in the language.
  2. Transcriptions in IPA and native script(s).
  3. Translations in English (which also act as prompts for the translation sentences).
  4. Detailed speaker metadata, including their demographic, educational and linguistic profile.
  5. Prompt in English and Hindi.

The full dataset contains the following -

  1. Translations of a minimum of 1000 carefully selected sentences. These sentences are selected to represent diverse morphosyntactic categories generally found in Indian languages such as demonstratives, classifiers, TAM morphology, and different sentence structures such as transitives and ditransitives. These sets of sentences are based on the standardised questionnaires built by Linguists for writing the sketch grammar of any human language, thereby, representing almost full range of morphosyntactic properties exhibited in the language. Unlike other benchmarks which focus largely on lexical level evaluation (aka domains), this is the first benchmark dataset that evaluates model's performance on a range of morphosyntactic structures, even the most uncommon ones.
  2. More than 600 narrative speeches collected across 8 domains. These are recorded by at least 2 speakers. Narrations, along with their prompts, can be used to evaluate AI models on a range of prompt-based tasks.
  3. Along with transcriptions in IPA and other scripts and translation, all data is interlinearly glossed at morphemic level. This gives a word-by-word meaning and morphosyntactic information, thereby, enabling evaluation of models on their deep grammatical knowledge, ability for cross-linguistic comparison and generalisation and reasoning capacity and skills in language-related puzzles. This allows for evaluating the model's capacity on reasoning tasks beyond mathematical reasoning tasks as well as their capacity to arrive at typological generalisations.
  4. In addition to the data and prompt itself, as mentioned earlier, we are also making available the detailed speaker metadata and prompt-level metadata viz domain, elicitation method, multilingual prompt, target grammatical category (for translation sentences), etc.
  5. In accordance with our data governance policy, all contributors to the dataset, including those recording the dataset and those transcribing it, are named and listed as data contributors.

Dataset Preparation

The speech included as part of this dataset was recorded by native speakers using a data collection application on a mobile device (Karya or in-house app, Atekho), by translating sentences from English or Hindi to the target language.

The recorded speech was validated and then transcribed manually by trained linguists, working with the native speakers, following the guidelines for the project using MATra Lab, a part of the LiFE Suite Ecosystem, developed by Unreal Tece LLP.

Project BoLI Guidelines

About Lambadi

Lambadi (ISO 639-3: lmn, Glottocode: lamb1269), also known as Banjari, Gor Boli, or Labani, belongs to the Indo-Aryan branch of the Indo-Iranian subfamily of the Indo-European language family. Within Indo-Aryan, it is traditionally classified under the Western Zone (Rajasthani group), showing strong historical ties to Gujarati and Rajasthani dialects, alongside regional lexical influences from the Dravidian languages.

Speakers

According to the 2011 Census of India, there are approximately 4.77 million speakers of Banjari/Lambadi (returned largely under the umbrella of Rajasthani/Hindi variants).

Distribution in India

The Lambadi-speaking nomadic and semi-nomadic Banjaras are widely dispersed across several Indian states. The principal concentrations of speakers are located in:

  • Telangana and Andhra Pradesh (where they are commonly referred to as Sugalis or Lambadis)
  • Karnataka (particularly northern and central districts)
  • Maharashtra
  • Rajasthan (their historical homeland)

Major Grammatical Features

  • Phonology: Lambadi possesses a typical Indo-Aryan phonological inventory, featuring a distinction between aspirated and unaspirated stops, retroflex consonants (/ʈ/, /ɖ/, /ɳ/, /ɭ/), and nasalized vowels. There is notable phonological variation influenced by surrounding Dravidian languages in southern states.
  • Morphology: Nominal morphology is characterized by two grammatical genders (masculine and feminine) and two numbers (singular and plural). Nouns inflect using postpositions rather than prepositions to mark cases (such as genitive -ro or -ko). Verbs agree with subjects in gender, number, and person.
  • Syntax: The basic word order is Subject-Object-Verb (SOV). Lambadi exhibits split ergativity, where transitive verbs in past/perfective tenses show ergative alignment, aligning the subject with an agentive case marker while the verb agrees with the object.

Further reading

About Project BoLI

Project BoLI is the flagship project of UnReaL-TecE LLP, which aims to build high-quality datasets for benchmarking and evaluating different kinds of AI tasks, including speech-to-text, machine translation, grammatical analysis and reasoning tasks and prompt-based evaluation of LLMs. While the project aims to build these datasets for every Indian language and variety, the primary focus is on over 1300 underserved languages and both first and second language varieties of major, scheduled languages. The project's uniqueness is not just limited to the kind of benchmarking tasks it supports (including proposing some novel tasks) and also the kind of languages and communities it supports but also in its contextualisation and implementation of a unique data governance model (not yet implemented anywhere across the globe), which mandates that all datasets released by the project and their derivatives (including the models) are co-owned by the community members and all contributors of the project, and any permission to use the dataset is not a transfer of ownership but a revocable license to use it. The conditions under which the license could be revoked are clearly mentioned as part of the BoLI License. The complete details of the project, languages and communities supported till now, the quantum of data available till now, its data governance model and other relevant documents are all publicly accessible at the project website.

Ethical Considerations, Consent, IPR and Attribution

Project BoLI and this repository represents our commitment to not only fair remuneration to the speakers of the language but also to co-ownership and equal IPR to all the contributors who have built the dataset. We believe this is the first step to move away from the extractive data collection and use practices and ensure fairness in our treatment of the community members. As such, we have listed all speakers and transcribers as Contributors to the dataset (we insist that they are co-owners of the dataset, even though the HuggingFace platform does not provide us an explicit way of stating that) and they are further recognised as Speakers and Annotators of the dataset. This dataset is only licensed to other researchers for use in their research projects. More details about licensing and commercial use conditions are given in the License and Commercial Use sections.

Project BoLI - Data Governance Policy

BoLI Ethics Principle & Pledge

Project BoLI - Digital Consent Form

Project BoLI - Field Recording of Oral Consent

Project BoLI - TnC

Dataset Access

Full dataset for the language can be browsed and queried on our app. The results of the evaluation of different models will also be made available on the same link. If you would like to access the full dataset for your own use or would like to work with us in collecting more data for any language or variety, or collaborate with us in this initiative in any other way, please get in touch with us.

Connected research

Relationships and public data

Record collections are summarized by type and remain available through the paginated public browser.

Contributors

Project collaborators

Tejavath Vinayak TagoreCommunity Collaborator
Neenavath Chetana LakshmiCommunity Collaborator
KORRA NITHINResearch Assistant

Reference

How to cite

A preferred citation has not been supplied.

Licence CC BY-NC-SA 4.0

Contact

Project contacts

  • account_circleProject ownerBoLI

For data access, use the request or workspace action in the project panel.