ModelRefs / Phi-4 by Microsoft — Benchmarks, Pricing & Review (2026)

Phi-4 by Microsoft — Benchmarks, Pricing & Review (2026)

Phi-4 (Microsoft): Phi-4 by Microsoft.. 16K context. Pricing: from $0.00007/1K in. Specs, benchmarks and code examples.

What this reference supports

Phi-4 is Microsoft's 14B small language model, trained with a synthetic-data-centric curriculum and released with open weights under an MIT license. It targets reasoning-heavy tasks at a size that fits single-GPU serving, making quality-per-parameter the central evaluation question against both larger open models and hosted small tiers.

Phi-4 is attributed to Microsoft in ModelRefs' canonical registry. Tracked modalities: Text input and output. Primary use cases considered on ModelRefs: Single-GPU reasoning, math, and analysis workloads on controlled infrastructure; Edge and cost-constrained deployments where a permissive license matters.

This ModelRefs profile is Provisional and pending review — decision-support material, not a final or universal ranking. Confirm current behavior, access, pricing, limits, licensing, and lifecycle in Microsoft's own documentation, and evaluate Phi-4 on representative workloads before implementation.

Benchmark & Evaluation

ModelRefs holds sourced benchmark evidence for Phi-4, but a benchmark score describes only its stated protocol and date — treat it as one input, not a guarantee of real-world performance, and evaluate Phi-4 on representative workloads before selecting it.

  • No benchmark score is imported into this editorial record. Provider-reported evaluations support scoped notes only; canonical score records are governed separately with their own provenance.
  • The technical report documents reported evaluations and the synthetic-data methodology; reproduction depends on harness and prompt configuration.

Implementation considerations

  • Evaluate on your actual task mix: synthetic-data-trained models can show uneven strengths between benchmark-style reasoning and messy real-world inputs.
  • Verify context-window and template requirements from the model card; small models are sensitive to prompt-format drift.
  • Open weights distributed via Hugging Face under MIT; also available through Azure AI model catalogs.
  • Azure-hosted and self-hosted deployments differ in versioning, quotas, and data controls.

Risks and limitations

  • Open-weight results depend on the exact runtime, precision, quantization, and prompt template; reference results do not transfer automatically.
  • The release-specific license and acceptable-use policy must be reviewed before commercial deployment.

Source coverage

This reference is Provisional. Model behavior, access, pricing, limits, and lifecycle can change; verify the linked provider documentation and run task-specific evaluations before implementation.

Known coverage gaps:

  • Independent evaluation reproduction and real-world (non-benchmark) task evidence is incomplete.
  • Azure catalog version mapping is not attached.

Sources

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Phi-4 by Microsoft — Benchmarks, Pricing & Review (2026).