ModelRefs / Latency Benchmarks — Top AI Models
Latency Benchmarks — Top AI Models
First-token latency, tokens/sec, streaming throughput. Quantify real-world responsiveness for interactive UX.
Overview
First-token latency, tokens/sec, streaming throughput.
What this category is for: Quantify real-world responsiveness for interactive UX.
Benchmarks in this category
- First-Token Latency — Time from request to first streamed token (lower is better).
- Output Tokens/Sec — Sustained streaming throughput once generation starts.
- Time-to-First-Token P95 — 95th-percentile first-token latency under 256-token prompt.
- Throughput @ Batch 32 — Sustained tokens/sec under batched server-side inference.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Latency Benchmarks — Top AI Models.