Discover free LLM API models across providers with real speed and latency benchmarks.Free models 2773Providers 157Free offerings 8727
| Model | Providers | Speed | Latency | Tests |
|---|---|---|---|---|
阿里巴巴Free Alibaba Qwen3.5 Omni Flash is a multimodal model in the Qwen series, offering text, image, video, and audio understanding capabilities. | +3 moreShow fewer | N/A | N/A | 0 |
Zhipu AIFree Zhipu AI GLM-4 FlashX is an accelerated GLM-4 Flash build for sub-second responses in customer support bots and high-QPS API gateways. | N/A | N/A | 0 | |
Google |
LMSpeed tracks 2773 LLM models available for free across 157 API providers. Free tiers vary by provider — some offer limited daily requests, others provide free credits for new users. All speed data is from real API tests.
Find a free model, compare its providers, then open the provider or model detail page before you start testing.
Use search, model-family filters, and capability tags to narrow the directory to the free models you actually want to try.
Check provider count, benchmark volume, throughput, and first-token latency so you are not choosing from price alone.
Review the provider page for availability, health checks, pricing notes, and any free-tier limits before building on it.
Common questions about free model tiers and how to use this directory.
| N/A |
| N/A |
| 0 |
MiniMax Voice Clone is a voice cloning model that replicates speaker characteristics from short reference audio, enabling personalized TTS for dubbing, assistants, and creative media production. | N/A | N/A | 0 |
Alibaba Qwen3.5 Omni Plus is the flagship omnimodal model in the Qwen series, natively understanding text, images, audio, and video with a 256K context window, real-time speech interaction, and agentic tool use. | N/A | N/A | 0 |
DeepSeek V3.2 Reasoner is a reasoning model in the DeepSeek series, designed for complex reasoning, problem-solving, and analytical tasks. | N/A | N/A | 0 |
MiniMax M2.1 Lightning is an ultra-fast MiniMax tier for real-time voice companions, streaming chat, and latency-sensitive multimodal experiences. | N/A | N/A | 0 |
Google Gemini 2.5 Flash Live is a realtime audio model in the Gemini series, supporting low-latency speech and conversational interactions. | N/A | N/A | 0 |
OpenAI GPT-4o DALL-E is an image generation model combining GPT-4o reasoning with DALL-E image synthesis, enabling detailed visual creation from conversational prompts. | N/A | N/A | 0 |
Zhipu AI GLM-4.1v Thinking Flash is a reasoning model in the GLM series, designed for complex reasoning, problem-solving, and analytical tasks. | N/A | N/A | 0 |
Google Gemini 2.5 Pro 1M is a high-capability language model in the Gemini series with a 1 million-token context window, offering enhanced reasoning, code generation, and multimodal capabilities. | N/A | N/A | 0 |
Alibaba Qwen3.6 Plus Thinking is a reasoning-focused variant in the Qwen series, designed for complex reasoning and problem-solving tasks. | N/A | N/A | 0 |
Google Gemini 3.1 Flash extends the Gemini 3 Flash line with improved reasoning and multimodal accuracy for production assistants and search-augmented apps. | N/A | N/A | 0 |
Zhipu AI GLM-4 AirX is an enhanced lightweight GLM-4 Air variant with higher throughput for enterprise chatbots and bilingual knowledge bases. | N/A | N/A | 0 |
Google Gemini Embedding is an embedding model, designed for generating vector representations of text for retrieval and semantic search. | N/A | N/A | 0 |
Google Gemini 2.0 Flash Live 001 is a realtime audio model in the Gemini series, supporting low-latency speech and conversational interactions. | N/A | N/A | 0 |
ReasoningToolsFilesOpen WeightsXiaomi MiMo-V2-Omni is the omnimodal model in the V2 series on the Xiaomi MiMo API platform, supporting text, image, video, and audio understanding within a unified architecture. Pricing: 1x token consumption (baseline). | N/A | N/A | 0 |
OpenAI GPT-3.5 Net is a language model in the GPT-3.5 series, offering general-purpose reasoning, code generation, and multimodal capabilities. | N/A | N/A | 0 |
ReasoningToolsOpen Weights204.8KMiniMax M2.5 HighSpeed prioritizes token throughput and low latency for large-scale conversational AI, content generation, and API relay traffic. | N/A | N/A | 0 |
Zhipu AI GLM-4V Flash is a multimodal vision-language model in the GLM series, supporting both text and image understanding. | N/A | N/A | 0 |
Google Gemini Live 2.5 Flash is a realtime audio model in the Gemini series, supporting low-latency speech and conversational interactions. | N/A | N/A | 0 |
Arctic Embed LFree Arctic Embed L is an embedding model, designed for generating vector representations of text for retrieval and semantic search. | N/A | N/A | 0 |
FilesVision64KZhipu AI GLM-4.1v Thinking FlashX is a reasoning model in the GLM series, designed for complex reasoning, problem-solving, and analytical tasks. | N/A | N/A | 0 |
Alibaba Qwen3.5 Max is a high-capability language model in the Qwen series, offering enhanced reasoning, code generation, and multimodal capabilities. | N/A | N/A | 0 |
Alibaba Qwen3.5 Plus Thinking is a reasoning-focused variant in the Qwen series, designed for complex reasoning and problem-solving tasks. | N/A | N/A | 0 |
ReasoningToolsOpen Weights128KMicrosoft Phi 3.5 MoE Instruct is a mixture-of-experts instruction-tuned variant in the Phi series, optimized for following instructions and conversational tasks. | N/A | N/A | 0 |
Zhipu GLM-4.5-X is a premium GLM-4.5 variant optimized for complex agentic reasoning, tool orchestration, and higher-quality instruction following. | N/A | N/A | 0 |
Anthropic Claude Opus 4.6 Max is a high-capability language model in the Claude series, offering enhanced reasoning, code generation, and multimodal capabilities. | N/A | N/A | 0 |
Italia InstructFree Italia Instruct is an instruction-tuned language model, optimized for following instructions and conversational tasks. | N/A | N/A | 0 |
ChatMiniMax M2.1 HighSpeed is a performance-tuned MiniMax M2.1 build for high-concurrency chat, roleplay, and streaming text generation. | N/A | N/A | 0 |
ToolsOpen Weights262.1KMoonshot AI Kimi K2 Turbo is a fast Kimi K2 tier for long-context chat, document analysis, and Chinese-first assistants with competitive inference speed. | N/A | N/A | 0 |
128KZhipu AI GLM-4 Flash is a budget-friendly GLM-4 variant for everyday dialogue, lightweight coding help, and high-volume API consumption. | N/A | N/A | 0 |
DeepSeek Coder Instruct is a code-specialized variant in the DeepSeek series, optimized for code generation, debugging, and software development tasks. | N/A | N/A | 0 |
128KZhipu AI GLM-4 Air is a compact GLM-4 model for mobile assistants, on-device gateways, and cost-sensitive bilingual chat applications. | N/A | N/A | 0 |
Solar InstructFree Upstage Solar Instruct is an instruction-tuned language model in the Solar series, optimized for chat, summarization, and enterprise document workflows with strong Korean and English performance. | N/A | N/A | 0 |
Reasoning128KReasoningHuawei PanGu Pro MoE is a mixture-of-experts language model designed for enterprise Chinese-English tasks, reasoning, and scalable cloud inference. | N/A | N/A | 0 |
Zhipu GLM-4 Plus is a high-end GLM-4 tier optimized for complex dialogue, analysis, and enterprise assistant scenarios. | N/A | N/A | 0 |
ReasoningToolsFilesOpen WeightsZhipu AI GLM-4.6V Flash is a multimodal vision-language model in the GLM series, supporting both text and image understanding. | N/A | N/A | 0 |
TeleSpeechASRFree TeleSpeechASR is a speech-to-text model designed for accurate audio transcription and recognition, supporting voice input pipelines, meeting notes, and accessibility applications. | N/A | N/A | 0 |
Meta Llama 3.2 Nemoretriever 300m Embed v1 is an embedding model, designed for generating vector representations of text for retrieval and semantic search. | N/A | N/A | 0 |
128KReasoningToolsOpen WeightsMicrosoft Phi 4 Multimodal Instruct is a multimodal instruction-tuned variant in the Phi series, optimized for following instructions and conversational tasks. | N/A | N/A | 0 |
Tools文件Multimodal66KStepFun Step3 is a general-purpose language model with strong reasoning and instruction-following capabilities, designed for chat, coding assistance, and agentic workflows on the StepFun platform. | N/A | N/A | 0 |
AI21 Jamba 1.5 Large Instruct is a large instruction-tuned language model, optimized for following instructions and conversational tasks. | N/A | N/A | 0 |
Zhipu AI GLM-Z1 Air is a lightweight reasoning-oriented GLM-Z1 variant for math, logic puzzles, and structured problem solving at low cost. | N/A | N/A | 0 |
ReasoningToolsOpen Weights131.1KZhipu AI GLM-4.5 Flash is a speed-first GLM-4.5 model for interactive chat, function calling, and bilingual customer-service automation. | N/A | N/A | 0 |
VideoAsyncMiniMax Hailuo 02 is a video generation model that produces short video clips from text or image prompts, supporting creative content workflows on the MiniMax platform. | N/A | N/A | 0 |
Alibaba Qwen3.0 is an early Qwen3 generation model offering solid multilingual chat, instruction following, and cost-efficient API deployment. | N/A | N/A | 0 |
ReasoningTools131KReasoningRing Flash 2.0 is a fast and efficient language model, optimized for quick responses and high throughput. | N/A | N/A | 0 |
Open WeightsVision77ImageAn open-weight image generation model by Black Forest Labs in the FLUX series, designed for development and experimentation. | N/A | N/A | 0 |
ChatVisionAnthropic Claude 3 Opus was the top-tier Claude 3 model for complex analysis, long documents, and nuanced generation before later Opus revisions. | N/A | N/A | 0 |
ToolsFilesVisionAudioGoogle Gemini 1.5 Flash is a multimodal model with up to 1M tokens of context, ideal for long-document summarization, audio transcription, and vision tasks. | N/A | N/A | 0 |
EmbeddingBGE Large ZH V1.5 is an embedding model, designed for generating vector representations of text for retrieval and semantic search. | N/A | N/A | 0 |
Google Gemini 2.0 Pro is a high-capability language model in the Gemini series, offering enhanced reasoning, code generation, and multimodal capabilities. | N/A | N/A | 0 |
ReasoningToolsOpen Weights262.1KMoonshot AI Kimi K2 Thinking Turbo is a reasoning model in the Kimi series, designed for complex reasoning, problem-solving, and analytical tasks. | N/A | N/A | 0 |
Google Gemini 2.0 Flash Lite is a lightweight and cost-efficient language model in the Gemini series, optimized for fast responses at reduced cost. | N/A | N/A | 0 |
Google Gemini 2.0 Flash Lite 001 is a compact language model in the Gemini series, optimized for low-latency responses and efficient inference. | N/A | N/A | 0 |
ChatVisionAnthropic Claude 3 Sonnet balanced capability and efficiency across text and vision tasks in the Claude 3 family. | N/A | N/A | 0 |
ReasoningTools200KFilesAnthropic Claude 3.7 Sonnet improves reasoning depth and coding quality over earlier Sonnet releases while staying suitable for interactive products. | N/A | N/A | 0 |
OpenAI GPT-4 DALL-E is an image generation model integrated with GPT-4 capabilities, producing detailed images from natural language descriptions for design and content workflows. | N/A | N/A | 0 |
ToolsFilesVisionAudioGoogle Gemini 2.0 Flash introduces native tool use and agentic capabilities with fast multimodal inference for production chat and automation. | N/A | N/A | 0 |
A model–provider pair whose current input and output prices are both $0 per token. Some providers offer permanent free tiers, others only give one-time credits to new accounts. We mark an offering as free only while its public pricing is zero, and re-check pricing pages regularly.
Most have rate limits — requests per minute, daily caps, or context-length limits — and many require account signup with a verified phone or payment method. Some are time-limited promotions. Always read the provider's terms and quotas before depending on a free endpoint.
Some entries are community-run relays ("公益站") that bundle paid upstream keys and redistribute access for free at the operator's expense. They often advertise larger quotas and a broader model list than official free tiers, but reliability is much lower: operators can pull the plug or disappear overnight, pricing and quotas can change without notice, and many sites are invite-only — requiring a GitHub invite, a forum referral, or a closed community to register. Some keep signups disabled indefinitely. Treat them as best-effort backup channels; keep anything important on official paid endpoints.
Speed varies by model and provider. Sort the table by Most tested for the most reliable benchmarks, or pick a model family from the chips above to drill in. Each row shows median tokens per second and first-token latency from real API tests.
We send identical prompts to each provider through a five-round stress test, count output tokens with tiktoken, and measure both throughput (tokens per second) and time to first token. Numbers are aggregated as medians to resist outliers and refresh on a regular cadence.
For prototypes, side projects, and low-traffic tools, yes. Production traffic will usually hit a rate limit quickly. Treat the free tier as an evaluation channel: validate the model and provider, then move to a paid endpoint with the same model when you scale.
Either no provider currently offers it for free, the free promotion ended, or it has not been benchmarked yet. Open the model's main page to compare paid options, or let us know about a missing free provider via the feedback link in the footer.