ModelRefs / Output Tokens/Sec Leaderboard — AI Model Scores
Output Tokens/Sec Leaderboard — AI Model Scores
Sustained streaming throughput once generation starts. Current leaders, methodology, and citation sources for Output Tokens/Sec.
Overview
Sustained streaming throughput once generation starts.
How it is measured: Mean output tok/s over a 1024-token generation.
How this benchmark is scored
| Category | latency |
|---|---|
| Maximum score | 500 tok/s |
| Direction | Higher is better |
| Evidence depth | incomplete |
Primary source: https://artificialanalysis.ai/
Published results
| Model | Score |
|---|---|
| Llama 4 Scout | 175 |
| GPT-5 Mini | 142 |
| Mistral Large 2 | 92 |
| GPT-5 | 84 |
Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Output Tokens/Sec Leaderboard — AI Model Scores.