ModelRefs / First-Token Latency Leaderboard — AI Model Scores
First-Token Latency Leaderboard — AI Model Scores
Time from request to first streamed token (lower is better). Current leaders, methodology, and citation sources for First-Token Latency.
Overview
Time from request to first streamed token (lower is better).
How it is measured: Median over 100 requests, 256-token prompt, US-East endpoint.
How this benchmark is scored
| Category | latency |
|---|---|
| Maximum score | 2000 ms |
| Direction | Lower is better |
| Evidence depth | incomplete |
Primary source: https://artificialanalysis.ai/
Published results
| Model | Score |
|---|---|
| Llama 4 Scout | 290 |
| GPT-5 Mini | 320 |
| Mistral Large 2 | 410 |
| GPT-5 | 540 |
Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to First-Token Latency Leaderboard — AI Model Scores.