SCOOT 2.0 Spoken LLM Benchmark Atlas

A living atlas of spoken LLM benchmarks

Unofficial and human-unverified: initially built by Claude Fable 5, prompted by Hung-yi Lee. GPT-6 checked the institution and publication metadata on 2026-09-06. Read the audit, corrections and remaining uncertainties · Disclosure.

Overview

What counts, and why this exists

Method

How papers are found, and how one is judged to be a benchmark

Taxonomy

How the benchmarks are classified

Timeline

When each benchmark hit arXiv

Bars count benchmarks by the month of their first arXiv posting (v1). Click a bar to filter the list below; click again to clear.

Who builds them

Where the benchmarks come from

Institutions

Countries & territories

Academic institutions, companies, or both

Lead institution's region, by year

In both lists: model builders that also publish benchmarks

How institutions were attributedmethod & limits
Where they were published

From preprint to proceedings

Venues

Which community they publish in

Status, by year of first arXiv posting

Preprint to proceedings, and what the claims rest on

How each paper’s venue was establishedmethod & limits
Results

Model × benchmark, by category

Rows are spoken LLMs, columns are benchmarks. Every number is a link: click it to open the paper it was copied from, and hover for the exact table it came from. Blanks mean not reported, not zero.

Survey

The accompanying overview paper

Daily crawl

Fresh from arXiv

An automated crawl scans new arXiv postings in eess.AS, cs.SD, cs.CL and cs.MM every day and flags anything that looks like a spoken-LLM benchmark. Flagged items are reviewed before they enter the catalogue above.

Contribute

Found an error? Missing a benchmark?

Changelog

Full catalogue

Every benchmark

All entries, newest first. One line each — click to open the full record.