Evaluation
Last Updated: August 25, 2026
Table of Contents
Overview
At test phase, the participants submit test stimuli (from a previously unreleased/blind test set) provided by the organisers and processed by their systems for evaluations leveraging crowdsourced subjective test and objective metrics. Objective scores do not contribute to a submission’s final ranking, which is based on subjective scores from the crowdsourced battery detailed below, plus a memory-bandwidth penalty in Track 2. Objective metrics may, however, serve as a qualification screen to allocate the limited full-battery evaluation slots if the number of entries exceeds the available budget (see Ranking).
Previously released open test set and blind test set (LRAC 2025) are available to participants for use during the development phase, though only LRAC 2.0 test stimuli will be submitted for evaluation.
The crowdsourced subjective battery targets transparency, robustness to mild noise and reverb, and intelligibility in clean speech through three types of tests:
MUSHRA-1S
A MUSHRA-style test with a single test stimulus (“1S”) plus low-anchor and reference files, assessing clean-speech quality across English (ENG), Spanish (SPA), and Mandarin Chinese (CMN), and scaling toward transparency.
DCR – Degradation Category Rating
Compares the test stimulus against the input reference file, assessing robustness to mild noise and reverberation and preservation of simultaneous talkers (English only).
DRT – Diagnostic Rhyme Test
Assesses speech intelligibility across English (ENG), Spanish (SPA), and Mandarin Chinese (CMN); for Mandarin Chinese, both consonant and tonal differences are evaluated.
Bitrate Caps
The evaluation spans five bitrate caps – {1, 6, 12, 24, 32} kbps. Each cap is a ceiling: a system’s constant rate for a given cap may sit anywhere up to and including the listed value, but must not exceed it (for example, 16 kbps satisfies the 32 kbps cap, whereas 48 kbps does not). A system already transparent at a lower cap may reuse that solution unchanged at any higher cap. As in LRAC 1.0, all operating points must be served by a single system (including a single decoder) under CBR, including a mixture of bitrates within a single inference run. See the Rules for the full technical constraints.
Subjective Battery
The final ranking is based on the crowdsourced subjective battery detailed below, with ranking weights summing to 100%. The same battery and weighting applies to both tracks. For Track 2, final ranking is determined after the memory-bandwidth penalty is applied.
English carries the largest quality weight and spans the full set of bitrate caps for clean-speech quality and robustness to real-world conditions. Multilingual clean speech is assessed only at the higher-rate caps ({12, 32} kbps), enabling a like-for-like comparison. Evaluation under real-world light noise and reverb is conducted for English only, to manage scope and cost. The simultaneous-talker condition measures preservation of overlapping speech, a common real-life scenario, and is assessed at the higher-rate caps ({12, 32} kbps). Intelligibility (DRT) is extended to multilingual coverage, and assessed at the ultralow bitrate (1 kbps) only, as the critical aspect for redundancy coding; Mandarin Chinese is included as a tonal language. Quality is reported per bitrate and per language (not only aggregated), so any language dependence can be traced across the bitrate caps.
Test Materials
Test materials will include curated test sets focused on quality assessment under clean and real-world conditions; see Evaluation Datasets for further details. Following the conclusion of the challenge, these test sets will be released publicly for the benefit of the research community. For crowdsourced assessment of intelligibility, previously reported approach along with publicly released stimuli link will be utilized.
Transparency
For a given bitrate cap, a submission is considered transparent when its mean MUSHRA-1S score on clean speech is not significantly below that of the hidden clean reference. Opus at 32 kbps is evaluated separately as a conventional-codec benchmark and shown on the board as a contextual reference for clean-speech scores.
Objective Metrics
Objective metrics will primarily be provided for informational purposes, though a qualification round (which may use objective and/or subjective scores) may be employed if the number of submissions per track exceeds the allocated budget (see Ranking). Otherwise, objective speech metrics do not contribute to the final ranking or aggregate score. Objective-vs-subjective correlations will be analyzed and reported following the conclusion of the challenge; see OBJ.2025 for a related analysis of objective metrics for neural codec evaluation.
Ranking
The overall ranking rests on the weighted subjective battery above. In Track 2, a memory-bandwidth penalty is additionally applied, so that at comparable quality and bitrate, solutions with lower memory-bandwidth demand rank higher; Track 1 is ranked on subjective scores alone.
The evaluation budget supports at least 7 full-battery runs per Track. If more entries are submitted, the organizers reserve the right to employ a qualification round (using objective metrics, crowdsourced listening tests, or both) to determine which submissions advance to full-battery evaluation (see Rules). Should such a qualification round be required, details will be posted in the Notices section of the Challenge website. Note that if fewer than 7 entries are received for one of the tracks, the remaining full-battery evaluation slots may be allocated to the other track. Winners are the top-ranked submission in each track. There are no financial prizes.