ModelRefs / First-Token Latency Leaderboard — AI Model Scores

First-Token Latency Leaderboard — AI Model Scores

Time from request to first streamed token (lower is better). Current leaders, methodology, and citation sources for First-Token Latency.

Overview

Time from request to first streamed token (lower is better).

How it is measured: Median over 100 requests, 256-token prompt, US-East endpoint.

How this benchmark is scored

Categorylatency
Maximum score2000 ms
DirectionLower is better
Evidence depthincomplete

Primary source: https://artificialanalysis.ai/

Published results

ModelScore
Llama 4 Scout290
GPT-5 Mini320
Mistral Large 2410
GPT-5540

Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to First-Token Latency Leaderboard — AI Model Scores.