ModelRefs / GPT-4.1 by OpenAI — Benchmarks, Pricing & Review (2026)
GPT-4.1 by OpenAI — Benchmarks, Pricing & Review (2026)
GPT-4.1 (OpenAI): GPT-4.1 by OpenAI.. 1M context. Pricing: from $0.00200/1K in. Specs, benchmarks and code examples.
What this reference supports
GPT-4.1 is an OpenAI API model in the GPT-4.1 family, designed for instruction-following, coding, tool use, and long-input applications. Implementation decisions should compare its non-reasoning behavior, latency, context use, and current lifecycle with newer OpenAI model families.
GPT-4.1 is attributed to OpenAI in ModelRefs' canonical registry. Tracked modalities: Text input and output, Image input. Primary use cases considered on ModelRefs: Repository and document analysis; Tool calling, structured extraction, and instruction-heavy workflows.
This ModelRefs profile is Provisional and pending review — decision-support material, not a final or universal ranking. Confirm current behavior, access, pricing, limits, licensing, and lifecycle in OpenAI's own documentation, and evaluate GPT-4.1 on representative workloads before implementation.
Benchmark & Evaluation
ModelRefs currently has partial, narrow benchmark coverage for GPT-4.1. Treat the available benchmark evidence as one input to the decision, not a guarantee that GPT-4.1 is the strongest option for your workload, and evaluate it on representative workloads before selecting it.
- Provider-reported benchmark results should be interpreted with methodology, dataset, prompting, tool, sampling, and recency limitations in mind.
- G.18 registers OpenAI's source-scoped simple-evals GPQA result for gpt-4.1-2025-04-14. It supports Partial eligibility only; broader benchmark coverage and independent reproduction remain incomplete.
Implementation considerations
- Use a dated model snapshot when behavior stability matters.
- Evaluate long-context retrieval and instruction adherence on representative inputs rather than assuming full-window reliability.
- Hosted through eligible OpenAI API endpoints.
- Endpoint features, quotas, and retention controls must be checked for the exact account and model snapshot.
Risks and limitations
- Outputs can be incorrect or unsuitable for the intended task; use task-specific evaluation, grounding, and human review where consequences are material.
- API availability, model aliases, rate limits, data controls, regions, and prices are mutable and differ by product channel.
Source coverage
This reference is Provisional. Model behavior, access, pricing, limits, and lifecycle can change; verify the linked provider documentation and run task-specific evaluations before implementation.
Known coverage gaps:
- Provider-reported evaluations lack independent reproduction and a registered canonical evaluation date.
- Current endpoint-specific data controls need periodic review.
Sources
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to GPT-4.1 by OpenAI — Benchmarks, Pricing & Review (2026).