Datasets

Last Updated: September 1, 2026

Table of Contents

  1. Training Datasets
    1. Speech Datasets
    2. Official Speech Curation
    3. Noise Datasets
    4. Room Impulse Response Datasets
  2. Evaluation Datasets
  3. References

Training Datasets

LRAC 2.0 builds directly on the training-data foundation established for LRAC 1.0. The speech, noise, and room impulse response (RIR) source datasets permitted in LRAC 1.0 remain permitted, and the official curated subset is carried forward unchanged.

For LRAC 2.0, only the speech collection is extended. We add Mozilla Common Voice v26 and AISHELL-3 (OpenSLR 93) to increase the amount of speech available for training and to broaden its language coverage. No new noise or RIR datasets are introduced.

After applying the official data-preparation pipeline, the curated subset contains approx. 1,200 hours of speech, roughly 70% more than the approx. 700 curated hours provided in LRAC 1.0. It covers 12 languages, compared with four in LRAC 1.0, and increases the share of non-English speech from below 25% to about 65% of the subset.

The curation process also guards against heavy gender imbalance in the selected subset; the resulting distribution is reported below.

All training, validation, and hyperparameter tuning must use only the officially designated source datasets. Participants are not required to use the official curated subset: the full, non-curated versions of the permitted datasets may also be used. The official LRAC 2.0 data-preparation pipeline prepares either collection and produces JSONL manifests and ESPnet/Kaldi-compatible data directories. Installation, dataset-access, and execution instructions are provided in the repository README.

Speech Datasets

The official LRAC 2.0 curated subset combines the LRAC 1.0 speech selection – drawn from LibriTTS LibriTTS.2019, VCTK VCTK.2019, EARS EARS.2024, Librivox from DNS5 DNS5.2024, MLS MLS.2020, and GLOBE V2 GLOBE.2024 – with selected material from two additional public datasets:

  • Mozilla Common Voice v26 CommonVoice.2020 adds community-contributed speech across a wider range of languages and speakers.
  • AISHELL-3 (OpenSLR 93) AISHELL3.2021 adds high-fidelity Mandarin speech from multiple native speakers.

The official data-preparation pipeline selects a subset of recordings from these sources for use by the baseline systems. This curated subset is provided as a reproducible reference, not as the only permitted form of the data. Participants may instead use the full, non-curated versions of any permitted source dataset.

Training Speech Datasets Distribution

Figure 1: Speech datasets permitted for model training in the LRAC 2.0 Challenge.

Official Speech Curation

The following procedure describes how the official curated subset is built for the LRAC 2.0 baseline. Participants may use their own curation and preprocessing procedures.

The curated LRAC 1.0 speech set is retained without modification in the official curated subset. Recordings from the three added datasets are assessed using the same principal measures used for LRAC 1.0, including estimated signal-to-noise ratio (SNR), reverberation, and bandwidth.

Where source metadata is unavailable, speaker gender is estimated to support analysis of the resulting distribution. Speaker identity and per-speaker duration are also considered to limit over-representation by individual speakers. The official pipeline excludes material reserved for evaluation; participants using non-curated data must also ensure that reserved evaluation recordings are excluded from training.

For the added datasets, some filtering thresholds are relaxed relative to LRAC 1.0. This allows the official curated subset to include substantially more speech from the newly represented languages, at the cost of admitting a broader range of recording conditions and some lower-quality real-world samples.

The official curated subset has the following high-level characteristics:

  • Approximately 1,200 hours of speech, compared with approximately 700 curated hours in LRAC 1.0.
  • 12 languages, compared with four in LRAC 1.0.
  • A split of roughly 35% English / 65% non-English, compared with a non-English share below 25% in LRAC 1.0.
  • Audio resampled to 24 kHz by the official data-preparation pipeline. The source recordings are not uniformly sampled at 24 kHz, and participants are not required to use the provided resampler.
  • Representation of every evaluation language in the training data; LRAC 2.0 does not include a held-out evaluation language.

Training Speech Number of hours per language distribution

Figure 2: Curated training-speech duration by language.

Language coverage and total duration are the principal improvements in LRAC 2.0. Because the source datasets differ considerably in size, speaker composition, and available metadata, the distribution is not uniform across languages.

Training Speech Gender Distribution

Figure 3: Overall gender distribution of the curated training speech.

Although the gender ratios are not perfectly balanced, the overall distribution is not heavily skewed. It should still be read alongside the per-language results, as the availability and completeness of gender metadata differ between datasets and languages.

Training Speech Gender Distribution per language

Figure 4: Gender distribution of the curated training speech by language.

Noise Datasets

LRAC 2.0 permits the same noise source datasets as LRAC 1.0:

No additional noise datasets are introduced for LRAC 2.0. Participants may use either the official curated noise set or the full, non-curated versions of the permitted sources.

For the official reference set, the LRAC 1.0 noise curation removes silence and material dominated by speech or other human vocal sounds, while retaining a broad range of environmental, machine, and musical-instrument noise.

For the complete source list and details of the classification and curation process, see the LRAC 1.0 dataset documentation.

Room Impulse Response Datasets

LRAC 2.0 permits the same room impulse response datasets as LRAC 1.0:

No additional RIR datasets are introduced. Participants may use either the officially prepared collection or the full, non-curated versions of the permitted RIR datasets. These recordings represent a range of acoustic spaces and can be used to simulate reverberant conditions during model training.

For further details, see the LRAC 1.0 dataset documentation.

Back to top

Evaluation Datasets

Participants will have access to previously released open test set and blind test set (LRAC 2025) for use throughout the challenge runtime. A separate, previously unreleased (withheld) test set will be distributed during the test phase; it serves as the challenge’s generalization set and is used for the final assessment of clean and real-world speech. Intelligibility testing (DRT) uses previously released test sets described in SIT.2024.

Coverage across the evaluation battery:

  • Quality and robustness are assessed on English across the five bitrate caps, including preservation of simultaneous talkers at the higher-rate caps.
  • Clean-speech quality is additionally assessed for Spanish and Mandarin Chinese at the higher-rate caps.
  • Intelligibility is assessed on English, Spanish, and Mandarin Chinese at the lowest bitrate cap.

Final test-set contents, file counts, and release timing are to be announced. Provisional coverage is shown below.

Evaluation test sets by language

Figure 5: Evaluation test sets by language (file counts and release timing to be finalized).

Every evaluation language is represented in the permitted training sources. The open and blind test sets described above remain available to participants for development purposes throughout the challenge runtime, but this data is not permitted for training in the LRAC 2.0 Challenge.

Evaluation material comprises complete, curated sentences rather than segments cut in the middle of words. For noisy conditions, the background material is selected to avoid semantic overlap with the speech.

Following the conclusion of the challenge, the withheld test set will be released publicly for the benefit of the research community.

Back to top

References

  1. [LibriTTS.2019]
    • H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen and Y. Wu, “LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,” in Proc. Interspeech, Sept. 2019.
  2. [VCTK.2019]
    • J. Yamagishi, C. Veaux and K. MacDonald, “CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit (version 0.92),” The Centre for Speech Technology Research (CSTR), University of Edinburgh, 2019.
  3. [EARS.2024]
    • J. Richter, Y.-C. Wu, S. Krenn, S. Welker, B. Lay, S. Watanabe, A. Richard and T. Gerkmann, “EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation,” in Proc. Interspeech, Sept. 2024, pp. 4873-4877.
  4. [DNS5.2024]
    • H. Dubey, A. Aazami, V. Gopal, B. Naderi, S. Braun, R. Cutler, H. Gamper, M. Golestaneh and R. Aichner, “ICASSP 2023 Deep Noise Suppression Challenge,” IEEE Open J. Signal Process., vol. 5, pp. 725-737, 2024.
  5. [MLS.2020]
    • V. Pratap, Q. Xu, A. Sriram, G. Synnaeve and R. Collobert, “MLS: A Large-Scale Multilingual Dataset for Speech Research,” in Proc. Interspeech, Oct. 2020.
  6. [GLOBE.2024]
    • W. Wang, Y. Song and S. Jha, “GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech,” arXiv, June 2024, arXiv:2406.14875.
  7. [CommonVoice.2020]
    • R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers and G. Weber, “Common Voice: A Massively-Multilingual Speech Corpus,” in Proc. LREC, May 2020, pp. 4218-4222.
  8. [AISHELL3.2021]
    • Y. Shi, H. Bu, X. Xu, S. Zhang and M. Li, “AISHELL-3: A Multi-Speaker Mandarin TTS Corpus,” in Proc. Interspeech, Aug. 2021, pp. 2756-2760.
  9. [WHAM.2019]
    • G. Wichern, J. Antognini, M. Flynn, L. R. Zhu, E. McQuinn, D. Crow, E. Manilow and J. Le Roux, “WHAM!: Extending Speech Separation to Noisy Environments,” in Proc. Interspeech, Sept. 2019.
  10. [FSD50K.2022]
    • E. Fonseca, X. Favory, J. Pons, F. Font and X. Serra, “FSD50K: An Open Dataset of Human-Labeled Sound Events,” IEEE/ACM Trans. Audio, Speech, Lang. Process., vol. 30, pp. 829-852, 2022.
  11. [FMA.2017]
    • M. Defferrard, K. Benzi, P. Vandergheynst and X. Bresson, “FMA: A Dataset for Music Analysis,” in Proc. ISMIR, Oct. 2017.
  12. [Motus.2021]
    • G. Götz, S. J. Schlecht and V. Pulkki, “A dataset of higher-order Ambisonic room impulse responses and 3D models measured in a room with varying furniture,” in Proc. I3DA, Sept. 2021, pp. 1-8.
  13. [OpenSLR28.2017]
    • T. Ko, V. Peddinti, D. Povey, M. L. Seltzer and S. Khudanpur, “A study on data augmentation of reverberant speech for robust speech recognition,” in Proc. ICASSP, Mar. 2017, pp. 5220-5224.
  14. [SIT.2024]
    • L. Lechler and K. Wojcicki, “Crowdsourced Multilingual Speech Intelligibility Testing,” In Proc. ICASSP 2024, Seoul, South Korea, 2024, pp. 1441-1445.

Back to top