Discover free LLM API models across providers with real speed and latency benchmarks.Free models 2825Providers 158Free offerings 8616
| Model | Providers | Speed | Latency | Tests |
|---|---|---|---|---|
QwenFree Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview. It is a multimodal reasoning model intended for complex reasoning, visual understanding,... | N/A | N/A | 0 | |
GoogleFree ReasoningToolsFilesVisionGemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function... |
LMSpeed tracks 2825 LLM models available for free across 158 API providers. Free tiers vary by provider — some offer limited daily requests, others provide free credits for new users. All speed data is from real API tests.
Find a free model, compare its providers, then open the provider or model detail page before you start testing.
Use search, model-family filters, and capability tags to narrow the directory to the free models you actually want to try.
Check provider count, benchmark volume, throughput, and first-token latency so you are not choosing from price alone.
Review the provider page for availability, health checks, pricing notes, and any free-tier limits before building on it.
Common questions about free model tiers and how to use this directory.
| N/A |
| N/A |
| 0 |
gpt-5.5OpenAIFree ReasoningToolsFilesVision | +14 moreShow fewer | 40.54 t/s | 6.70 s | 105 |
claude-haiku-4-5-20251001-thinkingAnthropicFree 文本 | N/A | N/A | 0 |
claude-3-opus-latestClaudeFree 文本 | N/A | N/A | 0 |
claude-3-7-sonnet-20250219AnthropicFree ReasoningToolsFilesVision | +4 moreShow fewer | N/A | N/A | 0 |
claude-3-sonnet-20240229AnthropicFree ToolsFilesVision200K | +1 moreShow fewer | N/A | N/A | 0 |
PaddlePaddle/PaddleOCR-VL-1.5SiliconFlow (China)Free FilesOpen WeightsVision16.4K | N/A | N/A | 0 |
[free]DeepSeek-V4-FlashDeepSeekFree | N/A | N/A | 0 |
aisingapore/sea-lion-7b-instructFree | +1 moreShow fewer | N/A | N/A | 0 |
claude-3-opus-20240229AnthropicFree ToolsFilesVision200K | N/A | N/A | 0 |
claude-3-haiku-20241022ClaudeFree 文本 | N/A | N/A | 0 |
claude-3-5-haiku-latestClaudeFree 文本 | N/A | N/A | 0 |
claude-3-5-sonnet-latestClaudeFree 文本 | N/A | N/A | 0 |
Kwai-Kolors/KolorsFree | N/A | N/A | 0 |
Doubao-1.5-vision-pro-32k字节跳动Free 文本 | N/A | N/A | 0 |
Doubao-1.5-pro-256k字节跳动Free 文本 | N/A | N/A | 0 |
Cursor-c4ClaudeFree 文本 | N/A | N/A | 0 |
Qwen/Qwen2.5-Coder-7B-Instruct阿里巴巴Free | N/A | N/A | 0 |
[free]DeepSeek-V3.2DeepSeekFree | N/A | N/A | 0 |
abab5.5-chatMiniMaxFree 文本 | +1 moreShow fewer | N/A | N/A | 0 |
agnes-image-2.1-flashFree | N/A | N/A | 0 |
claude-3-haiku@20240307ClaudeFree 文本 | N/A | N/A | 0 |
claude-3-haiku-allClaudeFree 文本 | N/A | N/A | 0 |
claude-3-haiku-20240307AnthropicFree 文本 | +1 moreShow fewer | N/A | N/A | 0 |
claude-3-haiku-20240229ClaudeFree 文本 | N/A | N/A | 0 |
chatglm_std智谱Free 文本 | N/A | N/A | 0 |
claude-3-5-haiku-20241022AnthropicFree ToolsFilesVision200K | +3 moreShow fewer | N/A | N/A | 0 |
claude-3-5-sonnet-20241022AnthropicFree ToolsFilesVision200K | +4 moreShow fewer | N/A | N/A | 0 |
claude-3-5-sonnet-allClaudeFree 文本 | N/A | N/A | 0 |
7eapi-1Free | N/A | N/A | 0 |
FLUX.1-Kontext-proFree 生图 | N/A | N/A | 0 |
FunAudioLLM/SenseVoiceSmallFree | N/A | N/A | 0 |
K2.7Free | N/A | N/A | 0 |
Cursor-c4-thinkingClaudeFree 文本 | N/A | N/A | 0 |
Cursor-co4ClaudeFree 文本 | N/A | N/A | 0 |
Cursor-c37ClaudeFree 文本 | N/A | N/A | 0 |
Cursor-c37-thinkingClaudeFree 文本 | N/A | N/A | 0 |
01-ai/yi-large零一万物Free | N/A | N/A | 0 |
Cursor-co4-thinkingClaudeFree 文本 | N/A | N/A | 0 |
THUDM/GLM-4.1V-9B-Thinking智谱Free | N/A | N/A | 0 |
THUDM/GLM-Z1-9B-0414SiliconFlow (China)Free ReasoningTools131K | N/A | N/A | 0 |
[free]deepseek-v4-flashDeepSeekFree | N/A | N/A | 0 |
[free]deepseek-v4-proDeepSeekFree | N/A | N/A | 0 |
adept/fuyu-8bFree | N/A | N/A | 0 |
agnes-2.0-flashFree | N/A | N/A | 0 |
7eimage-2Free | N/A | N/A | 0 |
claude-3-7-sonnet-20250219-thinkingAnthropicFree 文本 | N/A | N/A | 0 |
babbage-002Free 文本 | +1 moreShow fewer | N/A | N/A | 0 |
chatglm_lite智谱Free 文本 | N/A | N/A | 0 |
bigcode/starcoder2-15bFree | N/A | N/A | 0 |
big-pickleopencode zenFree ReasoningTools200K | N/A | N/A | 0 |
claude-3-7-sonnet-latestClaudeFree 文本 | N/A | N/A | 0 |
claude-3-7-sonnet-thinkingClaudeFree ReasoningToolsFilesVision | N/A | N/A | 0 |
Doubao-1.5-lite-32k字节跳动Free 文本 | N/A | N/A | 0 |
chatglm_pro智谱Free 文本 | N/A | N/A | 0 |
chatglm_turbo智谱Free 文本 | N/A | N/A | 0 |
claude-2ClaudeFree 文本 | N/A | N/A | 0 |
claude-3-5-sonnet-20240620AnthropicFree ToolsFilesVision200K | +2 moreShow fewer | N/A | N/A | 0 |
LongCat-2.0Free | N/A | N/A | 0 |
A model–provider pair whose current input and output prices are both $0 per token. Some providers offer permanent free tiers, others only give one-time credits to new accounts. We mark an offering as free only while its public pricing is zero, and re-check pricing pages regularly.
Most have rate limits — requests per minute, daily caps, or context-length limits — and many require account signup with a verified phone or payment method. Some are time-limited promotions. Always read the provider's terms and quotas before depending on a free endpoint.
Some entries are community-run relays ("公益站") that bundle paid upstream keys and redistribute access for free at the operator's expense. They often advertise larger quotas and a broader model list than official free tiers, but reliability is much lower: operators can pull the plug or disappear overnight, pricing and quotas can change without notice, and many sites are invite-only — requiring a GitHub invite, a forum referral, or a closed community to register. Some keep signups disabled indefinitely. Treat them as best-effort backup channels; keep anything important on official paid endpoints.
Speed varies by model and provider. Sort the table by Most tested for the most reliable benchmarks, or pick a model family from the chips above to drill in. Each row shows median tokens per second and first-token latency from real API tests.
We send identical prompts to each provider through a five-round stress test, count output tokens with tiktoken, and measure both throughput (tokens per second) and time to first token. Numbers are aggregated as medians to resist outliers and refresh on a regular cadence.
For prototypes, side projects, and low-traffic tools, yes. Production traffic will usually hit a rate limit quickly. Treat the free tier as an evaluation channel: validate the model and provider, then move to a paid endpoint with the same model when you scale.
Either no provider currently offers it for free, the free promotion ended, or it has not been benchmarked yet. Open the model's main page to compare paid options, or let us know about a missing free provider via the feedback link in the footer.
Explore the latest published LLMs tracked by LMSpeed, then open a model to compare providers, pricing, benchmarks, and free availability.