Free LLM API Models
Discover free LLM API models across providers with real speed and latency benchmarks.Free models 2817Providers 153Free offerings 8420
Latest LLM models
Explore the latest published LLMs tracked by LMSpeed, then open a model to compare providers, pricing, benchmarks, and free availability.
- Ling 3.0 Flash VLinclusionAI
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
Released Sep 10, 2026undefined =1 undefined other undefined providers}} - DeepSeek V4.1 FlashDeepSeek
DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the cost-efficient tier of the V4.1 family. DeepSeek reports that it exceeds V4 Pro on performance, speed, and task...
Released Sep 10, 2026undefined =1 undefined other undefined providers}}undefined other undefined free}} - Nex-N2.5-MiniNex AGI
Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...
Released Sep 8, 2026undefined =1 undefined other undefined providers}} - Nex-N2.5-ProNex AGI
Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file...
Released Sep 8, 2026undefined =1 undefined other undefined providers}} - GPT-6 AstraOpenAI
GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...
Released Sep 4, 2026undefined =1 undefined other undefined providers}}undefined other undefined free}} - Ling 3.0 Flash SanteinclusionAI
Ling 3.0 Flash Sante is a health and medicine-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for...
Released Sep 4, 2026undefined =1 undefined other undefined providers}}
Free model directory
Compare free models by provider count, speed, and first-token latency from real API tests.
| Model | Providers | Speed | Latency | Tests |
|---|---|---|---|---|
codestral-embedFree | N/A | N/A | 0 | |
cloudflare-glm-5.2智谱Free 文本 | N/A | N/A | 0 | |
cf41b0e0dd7d53e5Free | N/A | N/A | 0 | |
braveFree | N/A | N/A | 0 | |
Cursor-co4ClaudeFree 文本代码 | N/A | N/A | 0 | |
brave-deep-researchFree | N/A | N/A | 0 | |
chatglm_lite智谱Free 文本通用 | N/A | N/A | 0 | |
claude-3-7-sonnet-latestClaudeFree 文本代码 | N/A | N/A | 0 | |
aioFree | N/A | N/A | 0 | |
agnes-image-2.5-flashFree | N/A | N/A | 0 | |
aisingapore/sea-lion-7b-instructFree | N/A | N/A | 0 | |
agnes-3.0-flashFree | N/A | N/A | 0 | |
agnes-2.5-flashFree | N/A | N/A | 0 | |
babbage-002Free 文本 | N/A | N/A | 0 | |
Cursor-c4-thinkingClaudeFree 文本代码Reasoning | N/A | N/A | 0 | |
agnes-image-2.1-flashFree | N/A | N/A | 0 | |
agnes-video-2.5-flashFree | N/A | N/A | 0 | |
agnes-video-v2.0Free | N/A | N/A | 0 | |
MiniMax-H3-Context-IRMiniMaxFree 文本 | N/A | N/A | 0 | |
MiniCPM5-2BFree | N/A | N/A | 0 | |
big-pickleopencode zenFree ReasoningTools200K | N/A | N/A | 0 | |
bigcode/starcoder2-15bFree | N/A | N/A | 0 | |
Cursor-co4-thinkingClaudeFree 文本代码Reasoning | N/A | N/A | 0 | |
buzz-freeFree | N/A | N/A | 0 | |
Kwai-Kolors/KolorsKolorsFree | N/A | N/A | 0 | |
cc-glm-5-turbo智谱Free 文本 | N/A | N/A | 0 | |
Cursor-c4ClaudeFree 文本代码 | N/A | N/A | 0 | |
chatglm_pro智谱Free 文本通用 | N/A | N/A | 0 | |
chatglm_turbo智谱Free 文本通用 | N/A | N/A | 0 | |
claude-2ClaudeFree 文本 | N/A | N/A | 0 | |
[free]DeepSeek-V4-Flash-0731DeepSeekFree | N/A | N/A | 0 | |
adept/fuyu-8bFree | N/A | N/A | 0 | |
claude-3-5-sonnet-20241022AnthropicFree ToolsFilesVision200K | N/A | N/A | 0 | |
claude-3-5-sonnet-allClaudeFree 文本代码 | N/A | N/A | 0 | |
claude-3-7-sonnet-thinkingClaudeFree ReasoningToolsFilesVision | N/A | N/A | 0 | |
coding-glm-4.6-free智谱Free 文本 | N/A | N/A | 0 | |
abab5.5s-chatMiniMaxFree 文本 | N/A | N/A | 0 | |
claude-3-haiku-20240229ClaudeFree 文本智能体Tools | N/A | N/A | 0 | |
claude-3-haiku-allClaudeFree 文本 | N/A | N/A | 0 | |
claude-3-haiku@20240307ClaudeFree 文本智能体Tools | N/A | N/A | 0 | |
claude-3-opus-latestClaudeFree 文本 | N/A | N/A | 0 | |
claude-3-sonnet-20240229AnthropicFree ToolsFilesVision200K | N/A | N/A | 0 | |
Doubao-1.5-vision-pro-32k字节跳动Free 文本长上下文 | N/A | N/A | 0 | |
claude-haiku-4-5-20251001-thinkingAnthropicFree 文本代码Reasoning | N/A | N/A | 0 | |
claude-haiku-4-5-instantAnthropicFree | N/A | N/A | 0 | |
claude-opus-4-0ClaudeFree 文本代码Reasoning智能体 | N/A | N/A | 0 | |
claude-opus-4-1-thinking-allClaudeFree 文本代码Reasoning智能体 | N/A | N/A | 0 | |
[free]DeepSeek-V4-FlashDeepSeekFree | N/A | N/A | 0 | |
claude-opus-4-20250514-thinkingAnthropicFree 文本代码Reasoning智能体 | N/A | N/A | 0 | |
claude-opus-4-5-20251101-thinking302.AIFree ReasoningToolsFilesVision | N/A | N/A | 0 | |
abab6-chatMiniMaxFree 文本 | N/A | N/A | 0 | |
[free]DeepSeek-V3.2DeepSeekFree | N/A | N/A | 0 | |
claude-opus-4-7-thinkingAnthropicFree 文本代码Reasoning智能体 | N/A | N/A | 0 | |
claude-opus-4-8-thinkClaudeFree 文本智能体Tools | N/A | N/A | 0 | |
SparkDesk-v1.1讯飞Free 文本通用 | N/A | N/A | 0 | |
claude-sonnet-4-0ClaudeFree 文本代码Reasoning智能体 | N/A | N/A | 0 | |
Qwen3.8-Flash-Next阿里巴巴Free | N/A | N/A | 0 | |
claude-sonnet-4-20250514-thinkingAnthropicFree 文本代码Reasoning智能体 | N/A | N/A | 0 | |
Qwen/Qwen2.5-7B-InstructSiliconFlow (China)Free Tools33K | N/A | N/A | 0 | |
Doubao-1.5-pro-32k字节跳动Free 文本长上下文 | N/A | N/A | 0 |
LMSpeed tracks 2817 LLM models available for free across 153 API providers. Free tiers vary by provider — some offer limited daily requests, others provide free credits for new users. All speed data is from real API tests.
How to use
Find a free model, compare its providers, then open the provider or model detail page before you start testing.
- 1
Filter by model family
Use search, model-family filters, and capability tags to narrow the directory to the free models you actually want to try.
- 2
Compare speed and latency
Check provider count, benchmark volume, throughput, and first-token latency so you are not choosing from price alone.
- 3
Open the provider details
Review the provider page for availability, health checks, pricing notes, and any free-tier limits before building on it.
Free LLM API FAQ
Common questions about free model tiers and how to use this directory.
What counts as a free LLM API on LMSpeed?
A model–provider pair whose current input and output prices are both $0 per token. Some providers offer permanent free tiers, others only give one-time credits to new accounts. We mark an offering as free only while its public pricing is zero, and re-check pricing pages regularly.
Are these free APIs really free? Any catch?
Most have rate limits — requests per minute, daily caps, or context-length limits — and many require account signup with a verified phone or payment method. Some are time-limited promotions. Always read the provider's terms and quotas before depending on a free endpoint.
Are community-run relays and non-profit aggregators included? Any extra caveats?
Some entries are community-run relays ("公益站") that bundle paid upstream keys and redistribute access for free at the operator's expense. They often advertise larger quotas and a broader model list than official free tiers, but reliability is much lower: operators can pull the plug or disappear overnight, pricing and quotas can change without notice, and many sites are invite-only — requiring a GitHub invite, a forum referral, or a closed community to register. Some keep signups disabled indefinitely. Treat them as best-effort backup channels; keep anything important on official paid endpoints.
Which free LLM API is fastest?
Speed varies by model and provider. Sort the table by Most tested for the most reliable benchmarks, or pick a model family from the chips above to drill in. Each row shows median tokens per second and first-token latency from real API tests.
How does LMSpeed measure speed and latency?
We send identical prompts to each provider through a five-round stress test, count output tokens with tiktoken, and measure both throughput (tokens per second) and time to first token. Numbers are aggregated as medians to resist outliers and refresh on a regular cadence.
Can I use a free LLM API in production?
For prototypes, side projects, and low-traffic tools, yes. Production traffic will usually hit a rate limit quickly. Treat the free tier as an evaluation channel: validate the model and provider, then move to a paid endpoint with the same model when you scale.
Why don't I see a specific model in this list?
Either no provider currently offers it for free, the free promotion ended, or it has not been benchmarked yet. Open the model's main page to compare paid options, or let us know about a missing free provider via the feedback link in the footer.
