Discover free LLM API models across providers with real speed and latency benchmarks.Free models 2768Providers 155Free offerings 8686
| Model | Providers | Speed | Latency | Tests |
|---|---|---|---|---|
APEX-qwen-3.5-plus阿里巴巴Free | N/A | N/A | 0 | |
APEX-kimi-k2-thinkingMoonshotFree | N/A | N/A | 0 | |
APEX-kimi-k2.5MoonshotFree | N/A | N/A | 0 | |
LMSpeed tracks 2768 LLM models available for free across 155 API providers. Free tiers vary by provider — some offer limited daily requests, others provide free credits for new users. All speed data is from real API tests.
Find a free model, compare its providers, then open the provider or model detail page before you start testing.
Use search, model-family filters, and capability tags to narrow the directory to the free models you actually want to try.
Check provider count, benchmark volume, throughput, and first-token latency so you are not choosing from price alone.
Review the provider page for availability, health checks, pricing notes, and any free-tier limits before building on it.
Common questions about free model tiers and how to use this directory.
| N/A |
| N/A |
| 0 |
APEX-minimax-m2.5MinimaxFree | N/A | N/A | 0 |
TTSOpenAIFree | N/A | N/A | 0 |
qwen3 Vision Model阿里巴巴Free | N/A | N/A | 0 |
kimi-k2-thinking-agent(次模型)MoonshotFree | N/A | N/A | 0 |
APEX-gemini-2.5-flash-liteGoogleFree | N/A | N/A | 0 |
B-gemini-3.1-pro-preview-thinkingGoogleFree | N/A | N/A | 0 |
B-gemini-3.1-pro-preview-maxthinkingGoogleFree | N/A | N/A | 0 |
B-gemini-3.1-pro-previewGoogleFree | N/A | N/A | 0 |
APEX-deepseek-v3.2-chatDeepSeekFree | N/A | N/A | 0 |
APEX-deepseek-r1DeepSeekFree | N/A | N/A | 0 |
coder-model-KC阿里巴巴Free | N/A | N/A | 0 |
B-gemini-3-pro-preview-thinkingGoogleFree | N/A | N/A | 0 |
B-gemini-3-pro-previewGoogleFree | N/A | N/A | 0 |
B-gemini-3-flash-preview-thinkingGoogleFree | N/A | N/A | 0 |
APEX-claude-sonnet-4-thinkingAnthropicFree | N/A | N/A | 0 |
APEX-deepseek-v3DeepSeekFree | N/A | N/A | 0 |
APEX-gemini-2.5-flashGoogleFree | N/A | N/A | 0 |
APEX-deepseek-v3.1DeepSeekFree | N/A | N/A | 0 |
APEX-deepseek-v3.2DeepSeekFree | N/A | N/A | 0 |
B-gemini-3-flash-previewGoogleFree | N/A | N/A | 0 |
B-gemini-2.5-pro-thinkingGoogleFree | N/A | N/A | 0 |
B-gemini-2.5-proGoogleFree | N/A | N/A | 0 |
B-gemini-2.5-flashGoogleFree | N/A | N/A | 0 |
APEX-qwen3-vl-plus阿里巴巴Free | N/A | N/A | 0 |
Open WeightsAudio8.2K8KXiaomi MiMo-V2-TTS is a text-to-speech model in the MiMo series, optimized for natural speech synthesis and voice generation tasks. | N/A | N/A | 0 |
Anthropic Claude Opus 4.7 Max is a high-capability language model in the Claude series, offering enhanced reasoning, code generation, and multimodal capabilities. | N/A | N/A | 0 |
[L]gemini-3-pro-previewGoogleFree | N/A | N/A | 0 |
[L]gemini-3.1-flash-image-previewGoogleFree | N/A | N/A | 0 |
[L]gemini-3-flash-previewGoogleFree | N/A | N/A | 0 |
DeepSeek V4 is DeepSeek's next-generation foundation model family, built for advanced reasoning, coding, math, and long-context agent applications. | N/A | N/A | 0 |
Alibaba Qwen1.8B Long Context is a compact language model in the Qwen series, optimized for low-latency responses and efficient inference. | N/A | N/A | 0 |
Alibaba Qwen1.8B is a compact language model in the Qwen series, optimized for low-latency responses and efficient inference. | N/A | N/A | 0 |
Meta Llama 3.1 extends the Llama 3 family with stronger reasoning, tool use, and long-context support across 8B to 405B scales. | N/A | N/A | 0 |
ToolsOpen Weights128KMeta Llama 3.3 is an updated Llama 3 open model with improved instruction following, multilingual support, and efficient inference. | N/A | N/A | 0 |
Dracarys Llama 3.1 Instruct is an instruction-tuned variant, optimized for following instructions and conversational tasks. | N/A | N/A | 0 |
Usdcode Llama 3.1 Instruct is an instruction-tuned variant, optimized for following instructions and conversational tasks. | N/A | N/A | 0 |
DeepSeek R1 Distill Qwen 1.5B is a reasoning model in the DeepSeek series, designed for complex reasoning, problem-solving, and analytical tasks. | N/A | N/A | 0 |
Tools33KOpen Weights128KMeta Llama 3.1 Instruct is an instruction-tuned variant in the Llama series, optimized for following instructions and conversational tasks. | N/A | N/A | 0 |
DeepSeek V2.5 is a language model in the DeepSeek series, offering general-purpose reasoning, code generation, and multimodal capabilities. | N/A | N/A | 0 |
ReasoningTools128KChatDeepSeek V3.1 is an open-weights style frontier model from DeepSeek with strong math, coding, and Chinese-English bilingual reasoning. | N/A | N/A | 0 |
MiniMax File is the MiniMax API endpoint for retrieving and managing uploaded files and generated outputs such as video clips and synthesized speech assets. | N/A | N/A | 0 |
Alibaba Qwen3.5 Plus Image Edit is an image editing model in the Qwen 3.5 series, supporting instruction-based image modification, inpainting, and visual refinement tasks. | N/A | N/A | 0 |
Google Gemini 2.5 Pro DeepSearch is a search-augmented language model in the Gemini series, integrating web retrieval to provide up-to-date answers. | N/A | N/A | 0 |
DeepSeek V3 Turbo is a throughput-optimized DeepSeek V3 tier for large-scale chat, coding assistance, and MoE inference with strong price-to-performance. | N/A | N/A | 0 |
MiniMax Voice Design is a voice synthesis model that creates custom voices from stylistic prompts, enabling personalized TTS profiles for chatbots, narration, and interactive applications. | N/A | N/A | 0 |
ChatOpenAI GPT-4o Study is an instruction-tuned variant in the GPT-4 series, optimized for following instructions and conversational tasks. | N/A | N/A | 0 |
Google Gemini 1.5 Pro 002 is a high-capability language model in the Gemini series, offering enhanced reasoning, code generation, and multimodal capabilities. | N/A | N/A | 0 |
DeepSeek R1 Turbo is a reasoning model in the DeepSeek series, designed for complex reasoning, problem-solving, and analytical tasks. | N/A | N/A | 0 |
Google Gemini 2.5 Flash DeepSearch is a search-augmented language model in the Gemini series, integrating web retrieval to provide up-to-date answers. | N/A | N/A | 0 |
MiniMax File Upload is the MiniMax API file-management endpoint for uploading audio, video, and image assets used in voice cloning, video generation, and multimodal understanding workflows. | N/A | N/A | 0 |
文本Anthropic Claude 2.0 introduced stronger reasoning and safer long-form generation for early production Claude deployments. | N/A | N/A | 0 |
Alibaba Qwen3.5 Omni Flash is a multimodal model in the Qwen series, offering text, image, video, and audio understanding capabilities. | N/A | N/A | 0 |
Zhipu AI GLM-4 FlashX is an accelerated GLM-4 Flash build for sub-second responses in customer support bots and high-QPS API gateways. | N/A | N/A | 0 |
Google Gemini Robotics ER 1.5 is a robotics-focused multimodal model in the Gemini series, designed for embodied reasoning and physical task planning. | N/A | N/A | 0 |
MiniMax Voice Clone is a voice cloning model that replicates speaker characteristics from short reference audio, enabling personalized TTS for dubbing, assistants, and creative media production. | N/A | N/A | 0 |
Alibaba Qwen3.5 Omni Plus is the flagship omnimodal model in the Qwen series, natively understanding text, images, audio, and video with a 256K context window, real-time speech interaction, and agentic tool use. | N/A | N/A | 0 |
A model–provider pair whose current input and output prices are both $0 per token. Some providers offer permanent free tiers, others only give one-time credits to new accounts. We mark an offering as free only while its public pricing is zero, and re-check pricing pages regularly.
Most have rate limits — requests per minute, daily caps, or context-length limits — and many require account signup with a verified phone or payment method. Some are time-limited promotions. Always read the provider's terms and quotas before depending on a free endpoint.
Some entries are community-run relays ("公益站") that bundle paid upstream keys and redistribute access for free at the operator's expense. They often advertise larger quotas and a broader model list than official free tiers, but reliability is much lower: operators can pull the plug or disappear overnight, pricing and quotas can change without notice, and many sites are invite-only — requiring a GitHub invite, a forum referral, or a closed community to register. Some keep signups disabled indefinitely. Treat them as best-effort backup channels; keep anything important on official paid endpoints.
Speed varies by model and provider. Sort the table by Most tested for the most reliable benchmarks, or pick a model family from the chips above to drill in. Each row shows median tokens per second and first-token latency from real API tests.
We send identical prompts to each provider through a five-round stress test, count output tokens with tiktoken, and measure both throughput (tokens per second) and time to first token. Numbers are aggregated as medians to resist outliers and refresh on a regular cadence.
For prototypes, side projects, and low-traffic tools, yes. Production traffic will usually hit a rate limit quickly. Treat the free tier as an evaluation channel: validate the model and provider, then move to a paid endpoint with the same model when you scale.
Either no provider currently offers it for free, the free promotion ended, or it has not been benchmarked yet. Open the model's main page to compare paid options, or let us know about a missing free provider via the feedback link in the footer.
Explore the latest published LLMs tracked by LMSpeed, then open a model to compare providers, pricing, benchmarks, and free availability.
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...
Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...
Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...