A living atlas of spoken LLM benchmarks
Unofficial and human-unverified: initially built by Claude Fable 5, prompted by Hung-yi Lee. GPT-6 checked the institution and publication metadata on 2026-09-06. Read the audit, corrections and remaining uncertainties · Disclosure.
What counts, and why this exists
How papers are found, and how one is judged to be a benchmark
How the benchmarks are classified
When each benchmark hit arXiv
Bars count benchmarks by the month of their first arXiv posting (v1). Click a bar to filter the list below; click again to clear.
Where the benchmarks come from
Institutions
Countries & territories
Academic institutions, companies, or both
Lead institution's region, by year
In both lists: model builders that also publish benchmarks
How institutions were attributedmethod & limits
From preprint to proceedings
Venues
Which community they publish in
Status, by year of first arXiv posting
Preprint to proceedings, and what the claims rest on
How each paper’s venue was establishedmethod & limits
Model × benchmark, by category
Rows are spoken LLMs, columns are benchmarks. Every number is a link: click it to open the paper it was copied from, and hover for the exact table it came from. Blanks mean not reported, not zero.
The accompanying overview paper
Fresh from arXiv
An automated crawl scans new arXiv postings in eess.AS, cs.SD, cs.CL and cs.MM every day and flags anything that looks like a spoken-LLM benchmark. Flagged items are reviewed before they enter the catalogue above.
Found an error? Missing a benchmark?
Changelog
Every benchmark
All entries, newest first. One line each — click to open the full record.