Benchmark structure
Three phases, three explicit outputs
Each phase has its own task definition, output contract, evaluation metrics, and demo surface.
Phase 1 · Text output
Speaker-Aware Transcription
Phase 1 separates complete room transcription from clue-conditioned target transcription. Phase 1.3 extends the same text task with stable speaker-timbre descriptions.
Preparing Phase 1 demos...
Phase 2 · Audio output
Target Voice Extraction
Phase 2 reuses the speaker-clue vocabulary but changes the model target from text to audio. Timeline-preserved and speech-only outputs are evaluated separately.
Preparing Phase 2 demos...
Acoustic Event Grounding
Given mixed audio and a natural-language description of an audible event, return one start/end interval. The task schema, candidate boundaries, and evaluation protocol are still under review.
Evaluation contracts
Outputs are scored within their phase
No metric is shared across incompatible text, timeline, and waveform outputs.
Deterministic selectionEvery public demo is selected explicitly in versioned configuration.
Audio verificationReferenced audio, target stems, and output contracts are validated before export.
Public-safe metadataThe static bundle contains no private filesystem paths or raw source data.