ModelRefs / Llama 3.1 405B by Meta — Benchmarks, Pricing & Review (2026)
Llama 3.1 405B by Meta — Benchmarks, Pricing & Review (2026)
Llama 3.1 405B (Meta): Llama 3.1 405B by Meta.. 128K context. Pricing: from $0.00270/1K in. Specs, benchmarks and code examples.
What this reference supports
Llama 3.1 405B is Meta's largest Llama 3.1 text model, distributed as downloadable model artifacts under a release-specific license for self-managed or third-party-hosted deployment rather than a first-party hosted API, making infrastructure and serving architecture a central part of the decision.
Use this page to assess whether your infrastructure can support a model at this scale, compare it against smaller Llama 3.1 tiers and other open-weight models, and check which coding, multilingual, and reasoning benchmarks such as MMLU and HumanEval apply to this specific release and parameter count rather than a differently sized sibling.
Open-weight access does not mean unrestricted use or a complete managed service. Review the release-specific license terms, budget for distributed serving, quantization, and hardware decisions at this parameter count, and run your own safety and quality evaluation before deployment, since Meta's own release evaluations are provider-reported and not independently reproduced or audited by ModelRefs here at all.
Benchmark & Evaluation
ModelRefs currently has partial, narrow benchmark coverage for Llama 3.1 405B. Treat the available benchmark evidence as one input to the decision, not a guarantee that Llama 3.1 405B is the strongest option for your workload, and evaluate it on representative workloads before selecting it.
- Provider-reported benchmark results should be interpreted with methodology, dataset, prompting, tool, sampling, and recency limitations in mind.
- Meta's release and research paper contain provider-reported evaluations; no score is imported into this editorial record.
Implementation considerations
- Budget hardware, distributed serving, quantization, tokenizer, and prompt-template compatibility.
- Review the release-specific license and run safety, quality, latency, and cost evaluations on the selected runtime.
- Model artifacts are distributed for licensed deployment.
- Managed partners expose separate serving stacks, regions, controls, and commercial terms.
Risks and limitations
- Model artifacts do not provide a managed production service; operators own serving, security, monitoring, evaluation, and incident response.
- Quantization, prompt templates, runtime versions, hardware, and fine-tuning can materially change observed behavior.
Source coverage
This reference is Provisional. Model behavior, access, pricing, limits, and lifecycle can change; verify the linked provider documentation and run task-specific evaluations before implementation.
Known coverage gaps:
- Runtime-specific performance and total-cost evidence is incomplete.
- Release-specific safety artifacts are not attached at claim level.
Sources
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Llama 3.1 405B by Meta — Benchmarks, Pricing & Review (2026).