Tracks
Table of Contents
- Track 1 – Transparency Codecs
- Track 2 – Memory-Efficient Transparency Codecs
- Encourage Design Diversity
The challenge comprises two tracks detailed below. Both share the same constraints, data, and evaluation battery; Track 2 adds a hard parameter cap and a memory-bandwidth penalty on top of Track 1’s requirements. You may enter one or both tracks.
Track 1 – Transparency Codecs
Development of low-complexity, low-latency speech codecs designed to achieve perceptual transparency of input speech as bitrate increases, including under mild noise and reverberation conditions, and across multiple languages.
Goal: Minimize perceptual speech degradation while meeting the complexity, latency and bitrate constraints detailed under challenge Rules. Track 1 has no parameter cap and no memory-bandwidth penalty, and is ranked purely on perceptual quality and intelligibility from the crowdsourced Evaluation battery.
Focus Areas
-
Quality scaling toward transparency across increasing bitrate caps in clean speech
-
Robustness to mild real-world distortions, e.g., background noise, reverberation
-
Behavior across languages: English, Spanish, and Mandarin Chinese
Track 2 – Memory-Efficient Transparency Codecs
Everything in Track 1, with additional emphasis on memory efficiency, rewarding leaner solutions at comparable quality and bitrate.
Goal: Achieve perceptual transparency while meeting the same complexity, latency and bitrate constraints and subject to two additional conditions below, as detailed in challenge Rules:
-
a hard cap of ≤1.5M parameters as a firm eligibility constraint, and
-
a theoretical memory-bandwidth penalty applied to the ranking score
Track 2 is ranked based on perceptual quality and intelligibility from the crowdsourced Evaluation battery, with the memory-bandwidth penalty applied to the final score.
Focus Areas
-
Quality scaling toward transparency at low memory-bandwidth cost
-
Robustness to mild real-world distortions, e.g., background noise, reverberation
-
Behavior across languages: English, Spanish, and Mandarin Chinese
Encourage Design Diversity
We explicitly invite diversity beyond the Conv-Encoder + RVQ + Conv/GAN-Decoder pattern that dominated LRAC 1.0 – e.g., hybrid parametric + neural designs, classical-DSP front ends with neural quantizers/decoders, diffusion/flow decoders, scalar/lattice/product quantization, multi-rate architectures, and neural-augmented classical codecs.