Building Better Speech Data for Indic Languages

Research
Yash VyasKrishna Rupaakula

Four controlled-access datasets bring 3,845.5 hours of Hindi, Telugu, Tamil and Bengali speech to teams working on ASR training, evaluation and error analysis.

Short, clean clips can make a speech model look ready before it is. The Vistaar benchmark gives a concrete example: in the paper's Hindi results, IndicWhisper's word error rate is 7.6% on studio-read IndicTTS and 26.8% on spontaneous telephone speech from Gramvaani. That is a 19.2-point gap within the same language.

That gap is why we are introducing Numo Hindi, Numo Telugu, Numo Tamil and Numo Bengali. Together, our four releases contain 353,349 recordings and 3,845.5 hours of transcribed audio. Across the four repositories, we provide same-language depth, complete reference-transcript coverage, longer recordings, speaker groupings, demographic metadata and per-example validation signals for finding failures that a short-clip score can hide.

1. What We Are Releasing

ReleaseRecordingsHoursReleased speaker profiles
Numo Hindi101,4501,314.62,536
Numo Telugu132,4881,224.23,456
Numo Tamil61,783561.6753
Numo Bengali57,628745.1973
Total353,3493,845.57,718 (sum of released speaker profiles)

Each of our private Hugging Face repositories points its default configuration to data-v3/*.parquet. Every row contains full-resolution WAV audio, a reference transcript, topic, duration and format details, a deterministic pseudonymous speaker_id, age, gender, country and region fields, validation scores and an audio checksum. Our cards use <license>; prospective users should agree on access and permitted-use terms with us before starting work.

The 7,718 total is the sum of the per-release profile counts (overlapped for multilingual speaker profiles) derived from each dataset’s pseudonymous speaker_id values. It does not imply that 7,718 distinct people appear across the combined collection. Within a release, teams can use those identifiers to construct speaker-disjoint splits.

2. Depth in Four Languages

Published hours answer a limited but useful question: how much labeled audio is available within one language for training, evaluation and error analysis? They do not measure transcript accuracy, demographic coverage, recording difficulty or licensing.

Figure 1. Published transcribed, validated or otherwise labeled hours by language, shown on a logarithmic scale. Sources: Kathbath/IndicSUPERB, Common Voice 26.0, FLEURS, IndicTTS, MUCS, Shrutilipi and Vaani. Dataset definitions and collection designs differ, so the bars compare published scale rather than quality. MUCS Bengali is Bengali-English code-switched.

Kathbath is the closest matched reference because its IndicSUPERB release table covers all four languages. It reports 606 hours across Hindi, Telugu, Tamil and Bengali. Our releases total 3,845.5 hours across the same set, about 6.3 times the labeled audio.

Larger national collections provide a different kind of coverage. Shrutilipi reports more than 6,400 hours of mined broadcast news across 12 languages. Vaani reports about 31,255 collected hours across 165 districts, including 2,043 transcribed hours. IndicVoices reports 23.7K hours across 22 languages, with 11.2K hours transcribed, but its current card does not publish a matched four-language breakdown. Shrutilipi's published per-language totals exceed ours in Hindi and Tamil, while Vaani's transcribed hours are smaller than ours in all four languages. Our scope is narrower by design: 561.6 to 1,314.6 transcribed hours in each released language.

3. Transcripts You Can Inspect

We ‌report ‌WER ‌and CER for each of the 353,349 rows in our Hindi, Telugu, Tamil and Bengali dataset shards. The average space-delimited transcript length is 114.3 tokens (Hindi), 62.3 (Telugu), 62.4 (Tamil) and 99.9 (Bengali).

The comparison corpora create their "reference" transcripts in very different ways. Common Voice has volunteers read prompted sentences and uses community voting to flag clips as validated, invalidated, or insufficiently reviewed. FLEURS is parallel read speech paired from FLORES-101 and offers both raw and normalized transcript fields. Kathbath also offers read speech but makes both clean and noisy transcript versions.

Shrutilipi approaches by aligning All India Radio audio to transcript PDFs accounting for OCR errors, extra text, and portions of untranscribed speech; their authors report 21 language reviewers evaluating their corpus. The Vaani paper describes current in the Vaani corpus describes prompt-image speech with an audio quality and transcription accuracy multi-stage automated and human review process; their dataset card states that 2,043 of the roughly 31,255 hours they collected have transcripts.

Our clearest counterexample to any claim that only Numo makes transcripts inspectable is IndicVoices, which states their maker-checker then superchecker process for producing raw and normalized transcripts and their current schema makes both available. Our dataset cards currently do not state the process by which Numo reference transcripts are produced, nor do they state the ASR validator and text normalization steps used to calculate WER and CER.

In our releases, the median WER is 5.56% (Hindi), 13.24% (Telugu), 20.34% (Tamil) and 9.40% (Bengali) and the median CER is 2.82%, 2.21%, 4.01%, 6.59% respectively. We intend these fields to be analysis tools enabling teams to rank examples, set explicit filters, and to report whether conclusions are sensitive to those filters. They are screening signals, not quality guarantees on the transcripts.

4. Demographic Coverage You Can Measure

4.1 Speaker Grouping and Experimental Control

Speaker ‌grouping ‌is ‌an important experimental control because it enables us to enforce speaker-disjoint training and evaluation sets, and to build speaker-aware regression slices.

Figure 2. Published speaker, contributor or released-profile counts by language. The units are related but not identical. Kathbath counts sum male and female speakers in the IndicSUPERB release table; Common Voice counts are contributors in the 26.0 release JSON.

Figure ‌2 ‌in ‌our releases shows distinct pseudonymous speaker_id values. Kathbath reports speakers and Common Voice reports contributors. Our profile counts outnumber Kathbath's speaker counts in all four languages, and our contributor counts in Hindi and Telugu, but Common Voice in Tamil and Bengali. We see these as measures of granularity, not evidence for demographic representation.

4.2 Gender and Age Distributions

We report gender metadata for all but four of our 7,718 release profiles. Overall the distribution is near-balanced at 3,842 male and 3,808 female profiles, with 68 profiles labeled other, prefer not to say or missing, and these aggregates mask large skews by language. Our Hindi release is 70.4% male, and our Telugu release is 65.2% female, while our Tamil release is near-balanced (47.1% male, 49.3% female) and our Bengali release is near-balanced (52.6% male, 47.2% female). We report these distributions as an audit of our existing metadata, not as evidence that our speaker population is demographically representative.

We capture age metadata for 7,714 profiles. Our participants span an age range from 18 to 76 in Hindi, 104 in Telugu, 63 in Tamil, and 70 in Bengali, with median ages 26, 28, 29, and 26 respectively. About 79% of our release profiles have speakers aged 18 to 34, and about 74 of our release profiles have speakers aged 55 or older. We treat the 104 age in Telugu as an outlier under active review. Numo datasets enable benchmarking by age, and our speaker distribution is notably younger than other corpora.

4.3 Comparison with Existing Corpora

Other corpora capture different subsets of these metadata dimensions. Kathbath reports counts by binary gender in its release table, without age. FLEURS reports gender and lacks age in its public schema. Common Voice 26.0 reports high-level demographics, but its official clip-level splits lack gender information for 47.0% of Hindi, 19.4% of Telugu, 65.6% of Tamil, and 22.6% of Bengali clips, and are similarly unpopulated by age. IndicVoices describes its sampling process in more detail than we currently do, with age tiers, gender, geography, occupation and demographic quotas, sampling 51K speakers. Vaani captures wider geographical coverage across districts and claims demographic diversity, but lacks matched distribution by language in its public dataset card.

Our releases have a different aim in this ecosystem: we match granular age, gender, region and topic metadata with longer audio, complete reference, full reference coverage and per-example validation scores for each of our four language targets. We can define evaluation cohorts, split speakers strictly, and verify whether we are seeing improvements in models for different ages, gender labels, geography, topic, length of recording or transcript filtering. Very few open datasets or benchmark datasets have this complete set of experimental controls matched per-language like we do.

5. Longer Recordings Expose Different Failures

Our recordings average 46.6 seconds in Hindi, 33.3 in Telugu, 32.7 in Tamil and 46.5 in Bengali. The official Common Voice 26.0 release reports averages of roughly 4 to 6 seconds for these languages. FLEURS averages, computed from its released file metadata, range from roughly 11 to 13 seconds.

Figure 3. Seconds per released recording or clip. FLEURS values are computed from released sample metadata; Common Voice values come from the official 26.0 release JSON.

These longer than average recordings provide the opportunity to investigate whether errors tend to accumulate later in an utterance, explore different segmentation approaches, and thoroughly exercise streaming or endpointing systems with inputs that more closely resemble real-world long-form voice interactions. These insights can be expressed as regression gates to ensure long request reliability in production by the team. Longer duration does not necessarily mean greater acoustic difficulty, but it provides a wider margin of operation for a speech model to lose context or exit decoding prematurely.

6. What the Per-Example Audit Fields Add

In ‌each ‌of ‌our dataset repositories, each row represents a single audio recording. We include per-example audit columns directly on each row so practitioners can perform pre-stratify pre-scoring filters before running their downstream evaluation pipelines.

All examples across all four releases include the baseline WER and CER metrics as well as a synthetic_suspicion_score between 0 and 1, and our dataset cards for Hindi, Telugu, Tamil, and Bengali note that every example in the released dataset has a score lower than 0.6. We include an audio_sha256 cryptographic checksum on every row so users can cryptographically verify the raw audio bytes.

Across each of our dataset repositories, every row corresponds to a single audio recording. We attach per-example audit fields directly to each row, providing practitioners with granular metadata for pre-scoring stratification and filtering before running downstream evaluation pipelines.

Every example across our four releases includes the baseline WER and CER metrics alongside a synthetic_suspicion_score. This score spans from 0 to 1, and our dataset cards for Hindi, Telugu, Tamil, and Bengali document that every released sample scores below 0.6. We also attach an audio_sha256 checksum to every row to ensure cryptographic verification of the raw audio bytes.

Figure 4. Numeric audit fields attached to each released example. This is an availability count, not a quality score. Common Voice contributes up_votes and down_votes; a zero means the compared release does not publish comparable numeric per-example audit columns.

6.1 Sensitivity Testing and Pre-Deployment Gates

We include these audit columns so teams can perform reproducible sensitivity analyses on different thresholds of these metrics to understand the impact of filtering on sample yield, dialect coverage, error rates on benchmark sets, etc. These slices can be used as regression gates before deployment to identify exactly which audio characteristics are causing which models to fail in production candidates.

We emphasize that we treat these fields as screening signals rather than certificates, and as such, our current dataset cards are agnostic to the exact ASR validation model and text normalization pipeline used to produce WER and CER, the exact synthetic detector architecture, versions, calibration distributions, and thresholding. We encourage researchers to publish their exact filtering thresholds and interpret performance differences across splits as sensitivity benchmarks.

We provide these audit columns to enable reproducible sensitivity analyses. Teams can evaluate a model across the full corpus, sweep specific WER, CER, or synthetic-suspicion thresholds, and measure how filtering alters sample yield, dialect coverage, and benchmark error rates. These structured slices can serve as pre-deployment regression gates, pinpointing exactly which audio characteristics trigger model failures in production candidates.

We treat these fields strictly as screening signals rather than absolute certificates of transcript accuracy or acoustic fidelity. Our current dataset cards use an in-house language-specific text-normalization pipeline and ASR models to derive WER and CER. We also use an in-built synthetic detector model to derive synthetic scores. We encourage researchers to publish their exact filtering thresholds and interpret performance variations across splits as sensitivity benchmarks.

7. From Research Question to Release Gate

Our ‌releases provide per-language depth, along with the controls to define cohorts by transcript filter, duration, speaker count, age, gender, region, and topic. Researchers can validate whether findings are robust to changes in the defined cohorts. Product managers can make the same evidential case to swap out one aggregate WER for a release decision based on named cohorts and failure modes.

We are launching Numo Hindi, Telugu, Tamil and Bengali under controlled access to teams working on building and evaluating ASR in these languages. Please take a look at the corresponding repositories on Hugging Face and reach out to us to have a conversation about access and terms of permitted use.