Phi 4 API Benchmarks, Pricing & Provider Data
Compare Phi 4 with another model
Choose a model to open its comparison page.
Phi 4 benchmark, API pricing, and provider data cover 2 API providers, with prices starting at $0.010/request. The page also shows measured API speed and first-token latency.
Microsoft Phi-4 is a compact open-weight language model in the Phi series, optimized for reasoning, math, code generation, and efficient on-device or API deployment.
Category Performance
Observed-capability estimates with uncertainty reported separately from independent benchmark families.
- Coverage
- 6 / 8
- Methodology
- V3.0
#1Math46.9Estimated2/4 Measured dimensions
#2Coding39.1Provisional1/4 Measured dimensions
#2Reasoning39.1RatedGlobal rank #633/4 Measured dimensions
#4Instruction following37.8Provisional1/4 Measured dimensions
#5Knowledge37.5Provisional1/4 Measured dimensions
#6Agents28.4Provisional1/4 Measured dimensions
Specifications
Input and output token limits for this model, plus how it ranks on long-context understanding.
Features
Rankings
Excels at
Falls behind in
Detailed scores
Updated: Sep 1, 2026Speed & latency
undefined metric} other undefined metrics}}
Output speed44.2 tok/s#66 / 77Time to first token0.91 s#24 / 77
Speed & latency
undefined metric} other undefined metrics}}
Pricing
undefined metric} other undefined metrics}}
Input price$0.125/M#17 / 181Output price$0.500/M#22 / 181
Pricing
undefined metric} other undefined metrics}}
Agents
V3.0undefined metric} other undefined metrics}} · Provisional
Score28.480% interval12.3–44.41/4 Measured dimensions
Agents
V3.0undefined metric} other undefined metrics}} · Provisional
Coding
V3.0undefined metric} other undefined metrics}} · Provisional
Score39.180% interval25.2–53.01/4 Measured dimensionsLiveCodeBench23.1%#102 / 115SciCode26.0%#175 / 206
Coding
V3.0undefined metric} other undefined metrics}} · Provisional
Reasoning
V3.0undefined metric} other undefined metrics}} · Rated
Score39.1#6380% interval30.4–47.73/4 Measured dimensionsMMLU-Pro71.4%#98 / 129GPQA57.5%#174 / 213
Reasoning
V3.0undefined metric} other undefined metrics}} · Rated
Knowledge
V3.0undefined metric} other undefined metrics}} · Provisional
Score37.580% interval21.5–53.51/4 Measured dimensionsArtificial Analysis Intelligence Index4.5#109 / 111AA-GPQA Diamond57.5#100 / 108
Knowledge
V3.0undefined metric} other undefined metrics}} · Provisional
Math
V3.0undefined metric} other undefined metrics}} · Estimated
Score46.980% interval35.1–58.72/4 Measured dimensionsAIME14.3%#52 / 68MATH-50081.0%#48 / 73
Math
V3.0undefined metric} other undefined metrics}} · Estimated
Multilingual
V3.0undefined metric} other undefined metrics}} · No data
No data80% interval30.8–69.20/4 Measured dimensions
Multilingual
V3.0undefined metric} other undefined metrics}} · No data
Multimodal
V3.0undefined metric} other undefined metrics}} · No data
No data80% interval30.8–69.20/4 Measured dimensions
Multimodal
V3.0undefined metric} other undefined metrics}} · No data
Instruction following
V3.0undefined metric} other undefined metrics}} · Provisional
Score37.880% interval21.8–53.81/4 Measured dimensionsAA-IFBench23.5#83 / 84
Instruction following
V3.0undefined metric} other undefined metrics}} · Provisional
OpenRouter endpoints
1 endpointsThird-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.
| Provider endpoint | Input | Output | 1d uptime | 30m latency | 30m throughput | Context / output |
|---|---|---|---|---|---|---|
DeepInfra deepinfra/bf16 | $0.070/M | $0.140/M | 100% | — | — | undefined tokens / undefined tokens |
Pricing Comparison
Compare Phi 4 API pricing across 2 providers. Prices range from $0.010/request to $57.53/M. 素墨API offers the lowest rate at $0.010/request.
Alternatives & Similar Models
GPT-5.4
gpt-5-4
OpenAI GPT-5.4 extends the GPT-5 family with stronger instruction following, deeper tool use, and improved performance on coding, math, and long-document analysis.
GPT-5.2
gpt-5-2
OpenAI GPT-5.2 is a GPT-5 series model emphasizing advanced reasoning, multimodal understanding, and high-quality outputs for complex enterprise workloads.
Claude Sonnet 4.6
claude-sonnet-4-6
Anthropic Claude Sonnet 4.6 extends the Sonnet line with improved tool use, coding reliability, and long-context performance for everyday production workloads.
Kimi K2.5
kimi-k2-5
Moonshot Kimi K2.5 is an open-weight multimodal agent model with native vision and text input, strong coding performance, and a 256K context window.
GLM-5
glm-5
Zhipu GLM-5 is Zhipu flagship GLM series model with enhanced reasoning, agent capabilities, and strong performance on Chinese enterprise and coding scenarios.
DeepSeek R1
deepseek-r1
DeepSeek R1 is a reasoning-focused language model in the DeepSeek series, designed for complex reasoning, problem-solving, and analytical tasks.
Frequently Asked Questions
- What benchmark data does Phi 4 include?
- LMSpeed shows Phi 4 benchmark context, API price, output speed, first-token latency, and provider data across 2 providers when those signals are available.
- What is the Phi 4 API price?
- Phi 4 has pricing from 2 providers, ranging from $0.010/request to $57.53/M. 素墨API has the lowest listed price.
- What does the Phi 4 API pricing table include?
- The Phi 4 API pricing table compares 2 providers by input price, output price, free tier status, speed, first-token latency, and recent health data when available.
- Which provider has the cheapest Phi 4 API pricing?
- 素墨API currently has the lowest listed Phi 4 price at $0.010/request across 2 providers.
- Can I compare Phi 4 API price and speed together?
- Yes. LMSpeed shows Phi 4 API price, output speed, first-token latency, and provider health on the same page so you can compare cost and performance together.
- Is Phi 4 API free?
- Phi 4 does not currently have a free API tier on LMSpeed. All 2 providers charge per token.
