Discover free LLM API models across providers with real speed and latency benchmarks.Free models 2775Providers 156Free offerings 8772
| Model | Providers | Speed | Latency | Tests |
|---|---|---|---|---|
Z.aiFree ReasoningToolsFilesVisionZhipu AI GLM-5V Turbo is a multimodal vision-language model in the GLM series, supporting both text and image understanding. | +2 moreShow fewer | N/A | N/A | 0 |
xAIFree Grok 4.20 Multi-Agent is a variant of xAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information... |
LMSpeed tracks 2775 LLM models available for free across 156 API providers. Free tiers vary by provider — some offer limited daily requests, others provide free credits for new users. All speed data is from real API tests.
Find a free model, compare its providers, then open the provider or model detail page before you start testing.
Use search, model-family filters, and capability tags to narrow the directory to the free models you actually want to try.
Check provider count, benchmark volume, throughput, and first-token latency so you are not choosing from price alone.
Review the provider page for availability, health checks, pricing notes, and any free-tier limits before building on it.
Common questions about free model tiers and how to use this directory.
| N/A |
| N/A |
| 0 |
ReasoningToolsFilesVisionxAI Grok 4.20 is a Grok 4 series model with real-time knowledge integration, strong reasoning, and conversational capabilities for search-augmented chat. | 88.54 t/s | 4.93 s | 15 |
KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integration. It builds on the agentic coding strengths of earlier versions,... | N/A | N/A | 0 |
ReasoningToolsFiles204.8KMiniMax M2.7 is a high-tier M2-series model tuned for complex reasoning, long-context dialogue, and production-grade API workloads. | Dext APIApiToken OnlineZero API +46 moreShow fewer初叶🍂Furry APIWSocket AI91VIP APIMoyanjdc APIFuture Hub霸气公益平台KternaDMXAPIAI Claw APIEnenCloud API933999 APILLM API情酱的API站CxyKevin API老魔公益站Hizui API百万APIAIO通用智能服务平台PICO AI42公益站DeepRouterKoru API随时跑路公益站Neb 公益站Koyeb AI GatewaySWT-APIZetaTechs APISmart APIOptAITommyLam API毫秒APIDibin84 API HubSeamee APISoul 公益站wuer的api站Mars HKMineWuer API左大臣星见雅 API6i2AI APIWAADRIMyNav AIC85 API温云AWA1 API | 43.91 t/s | 11.56 s | 10 |
ReasoningToolsFilesVisionOpenAI GPT-5.4 Nano is a compact language model in the GPT-5 series, optimized for low-latency responses and efficient inference. | N/A | N/A | 0 |
FastToolsReasoningFilesOpenAI GPT-5.4 Mini is a compact language model in the GPT-5 series, optimized for quick responses and high throughput. | Dext API兔子APIMoyanjdc API +45 moreShow fewerFuture Hubllm-2-apiDMXAPIVenlacyLiuwang API933999 APIFeng Love APICan API17NAS API天宫造物Codex EasyMagicAIMitchll-API艾可APIAIO通用智能服务平台Huainova 公益站汪汪中转站Liunew APIAIsa42公益站Codex APIDeepRouter随时跑路公益站Neb 公益站ZetaTechs APISmart APISwifllyLLMKJK APITommyLam API毫秒APIMentoe APIDibin84 API HubQQ CodeJuCode逆龙傲公益站Mars HKWzjself APIwzjself中转站左大臣AI APIDNSHEWAADRIMyNav AI至强APIGOU API | 144.48 t/s | 4.78 s | 40 |
ReasoningTools200KUsZhipu AI GLM-5 Turbo is a high-throughput GLM-5 tier for scalable inference, long-context reasoning, and production agents requiring strong Chinese-English performance. | N/A | N/A | 0 |
Seed-2.0-Lite is a versatile, cost‑efficient enterprise workhorse that delivers strong multimodal and agent capabilities while offering noticeably lower latency, making it a practical default choice for most production workloads across... | N/A | N/A | 0 |
ReasoningToolsFilesOpen WeightsAlibaba Qwen3.5 is a Qwen3 generation model with improved reasoning, multilingual support, and efficient inference for chat, coding, and agent applications. | 55.82 t/s | 10.52 s | 20 |
ReasoningToolsFilesVisionOpenAI GPT-5.4 Pro is a high-capability language model in the GPT-5 series, offering enhanced reasoning, code generation, and multimodal capabilities. | N/A | N/A | 0 |
CodexReasoningToolsFilesOpenAI GPT-5.4 extends the GPT-5 family with stronger instruction following, deeper tool use, and improved performance on coding, math, and long-document analysis. | Dext API兔子API180txt API +61 moreShow fewerMoyanjdc APIFuture Hubllm-2-apiKternaDMXAPICodeXEPrivnodeSkyAIVenlacyLiuwang API933999 APIFeng Love APICan APIHotaruAPI17NAS API情酱的API站天宫造物Codex Easy丸美小沐写作丸美小沐MagicAIMitchll-API艾可APIAIO通用智能服务平台ZenScale AI汪汪中转站Xiao WanLiunew APIAIsa42公益站Codex APIDeepRouter随时跑路公益站ZetaTechs APISmart APISwifllyLLMKJK APITommyLam API毫秒APIMentoe APIDibin84 API HubQQ CodeSeamee APIRinkoAICHB APIJuCodeHank Workspace APISoul 公益站逆龙傲公益站不知道叫啥Mars HKWzjself APIwzjself中转站6i2AI APIDNSHEWAADRIMyNav AI至强APIGOU APIOpenOpen8 API | 48.77 t/s | 6.54 s | 260 |
ReasoningToolsFilesVisionOpenAI GPT-5.3 is a GPT-5 series checkpoint tuned for stronger reasoning, coding, and agentic workflows on complex user tasks. | N/A | N/A | 0 |
Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across... | N/A | N/A | 0 |
Seed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference deployment. It delivers performance comparable to ByteDance-Seed-1.6, supports 256k context, four reasoning effort modes (minimal/low/medium/high), multimodal understanding,... | N/A | N/A | 0 |
ReasoningToolsFilesVisionAlibaba Qwen3.5 Flash is a fast multimodal language model in the Qwen series, delivering cost-efficient text, image, and video understanding with a 1M-token context window. | N/A | N/A | 0 |
CodexReasoningTools128KOpenAI GPT-5.3 Codex is a code-specialized variant in the GPT-5 series, optimized for code generation, debugging, and software development tasks. | 兔子APIMoyanjdc APIFuture Hub +49 moreShow fewerllm-2-apiKternaDMXAPIPrivnodeVenlacyLiuwang API933999 APIFeng Love API情酱的API站天宫造物Codex EasyMagicAIMitchll-APIAIO通用智能服务平台ZenScale AI汪汪中转站Liunew APIAIsa42公益站Codex APIDeepRouter随时跑路公益站ZetaTechs APISmart APISwifllyLLMKJK APITommyLam API毫秒APIMentoe APIDibin84 API HubQQ CodeSeamee APIRinkoAICHB APIJuCodeHank Workspace API逆龙傲公益站不知道叫啥Mars HKWzjself APIwzjself中转站AI APIDNSHEWAADRIMyNav AI至强API10dian-APIGOU APIOpenOpen8 API | 47.34 t/s | 4.16 s | 25 |
Aion-2.0 is a variant of DeepSeek V3.2 optimized for immersive roleplaying and storytelling. It is particularly strong at introducing tension, crises, and conflict into stories, making narratives feel more engaging.... | N/A | N/A | 0 |
Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation... | N/A | N/A | 0 |
ReasoningToolsFilesVisionGoogle Gemini 3.1 Pro is a Gemini 3 series model with advanced multimodal reasoning, long-context support, and strong performance on coding and analytical tasks. | 69.88 t/s | 18.61 s | 5 |
ReasoningToolsFilesVisionAnthropic Claude Sonnet 4.6 extends the Sonnet line with improved tool use, coding reliability, and long-context performance for everyday production workloads. | 兔子APIMoyanjdc APIKterna +43 moreShow fewerDMXAPIPrivnodeVenlacyCan API17NAS API情酱的API站天宫造物丸美小沐写作丸美小沐MagicAI艾可APIQWQ Chat APIAIO通用智能服务平台Fangyuan APILiunew APIAIsaDeepRouterKoyeb AI GatewaySWT-APIZetaTechs APISmart APIOptAISwifllyLLMTommyLam API毫秒APIMentoe APIDibin84 API HubQQ CodeNuizi APISeamee APIRinkoAICHB APIJuCodeHank Workspace API逆龙傲公益站不知道叫啥APDSMAI APIDNSHEWAADRIMyNav AI10dian-APIOpenOpen8 API | N/A | N/A | 0 |
ReasoningToolsOpen Weights204.8KMiniMax M2.5 is MiniMax's flagship text model for coding and agents, with SOTA-level programming and agentic performance, improved token efficiency, and fast high-TPS API deployment. | Dext API兔子APIZero API +38 moreShow fewerMoyanjdc APIFuture Hub霸气公益平台DMXAPIAI Claw API933999 API情酱的API站Mitchll-APIAIO通用智能服务平台Xiao WanAIsaDeepRouterKoru API随时跑路公益站Neb 公益站Koyeb AI GatewaySWT-APIZetaTechs APISmart APITommyLam API毫秒APISeamee APIRinkoAIwuer的api站Mars HKAPDSMMineWuer APIKFC API星见雅 API6i2MapleLeaf APIAI APISpaceshipDNSHEWAADRIC85 API温云OpenOpen8 API | 86.62 t/s | 6.48 s | 10 |
ReasoningTools200K204.8KZhipu GLM-5 is Zhipu flagship GLM series model with enhanced reasoning, agent capabilities, and strong performance on Chinese enterprise and coding scenarios. | 兔子APIZero APIMoyanjdc API +35 moreShow fewer霸气公益平台KternaDMXAPI933999 APIFeng Love APILLM API17NAS API情酱的API站Mitchll-API百万APIAIO通用智能服务平台AIsaPICO AIDeepRouterKoru API随时跑路公益站Neb 公益站Koyeb AI GatewaySWT-APIZetaTechs APISmart APIOptAITommyLam API毫秒APISeamee APIRinkoAICHB API不知道叫啥Mars HKMapleLeaf APIAI APISpaceshipWAADRIGOU APIAWA1 API | 52.98 t/s | 17.54 s | 15 |
Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By significantly scaling model capacity and reinforcement learning compute, it... | N/A | N/A | 0 |
ReasoningToolsFilesVisionAnthropic Claude Opus 4.6 is the most capable Claude Opus tier, optimized for complex analysis, long-horizon coding, and high-stakes enterprise reasoning workloads. | 兔子API我不是AI神Moyanjdc API +47 moreShow fewerKternaDMXAPICodeXEPrivnode丰思理 AISkyAIVenlacyCan APILLM API17NAS API情酱的API站天宫造物丸美小沐写作丸美小沐MagicAI艾可APIQWQ Chat APIAIO通用智能服务平台Xiao WanFangyuan APILiunew APIAIsa42公益站DeepRouter哈基米APIZetaTechs APIOptAISwifllyLLMTommyLam API毫秒APIMentoe APIDibin84 API HubQQ CodeNuizi APISeamee APIRinkoAICHB APIJuCodeHank Workspace APISoul 公益站逆龙傲公益站不知道叫啥AI APIWAADRIMyNav AI10dian-APIOpenOpen8 API | 43.29 t/s | 5.73 s | 40 |
ReasoningTools262KOpen WeightsStep 3.5 Flash is a fast and efficient language model, optimized for quick responses and high throughput. | N/A | N/A | 0 |
ReasoningToolsFilesOpen WeightsMoonshot Kimi K2.5 is an open-weight multimodal agent model with native vision and text input, strong coding performance, and a 256K context window. | Dext API兔子APIZero API +35 moreShow fewerMoyanjdc API霸气公益平台DMXAPIAI Claw API933999 APILLM API17NAS APIMitchll-APIAIO通用智能服务平台Real AI WANXiao WanAIsaE-larex's AI ProxyPICO AICodex APIDeepRouterKoru APINeb 公益站ZetaTechs API毫秒APISeamee APICHB APISoul 公益站wuer的api站APDSMMineWuer API星见雅 APIMapleLeaf APIAI APISpaceshipWAADRIMyNav AIC85 API温云OpenOpen8 API | 36.70 t/s | 26.01 s | 20 |
Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers exceptional performance while maintaining computational efficiency. Optimized... | N/A | N/A | 0 |
LFM 2.5 1.2B Instruct is a compact language model in the LFM series, optimized for low-latency responses and efficient inference. | N/A | N/A | 0 |
ReasoningToolsOpen Weights200KZhipu AI GLM-4.7 Flash is a speed-focused GLM variant for real-time chat, function calling, and bilingual enterprise copilots with competitive token economics. | 53.04 t/s | 20.54 s | 5 |
CodexReasoningToolsFilesOpenAI GPT-5.2 Codex is a code-specialized variant in the GPT-5 series, optimized for code generation, debugging, and software development tasks. | 30.05 t/s | 4.72 s | 5 |
Seed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding. It features a 256k context window and can generate outputs of... | N/A | N/A | 0 |
Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinking with a 256K context window. | N/A | N/A | 0 |
ReasoningToolsOpen Weights204.8KMiniMax M2.1 is an upgraded M2-series model with improved reasoning, multilingual chat, and agent capabilities for API-first production deployments. | 60.91 t/s | 6.39 s | 5 |
ReasoningTools200K205KZhipu GLM-4.7 is a flagship GLM release from Zhipu AI with advanced Chinese-English reasoning, coding, and agent features. | N/A | N/A | 0 |
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool... | N/A | N/A | 0 |
1000K金牌模型Multimodal长上下文Google Gemini 3 Flash is a next-generation fast multimodal model for responsive assistants, document understanding, and high-throughput API traffic. | 兔子API天絮 APIMoyanjdc API +36 moreShow fewerFuture HubElysiver APIDMXAPIPrivnode小辣椒情酱的API站丸美小沐写作丸美小沐艾可API百万APIAIO通用智能服务平台Xiao WanAIsa42公益站DeepRouterKoyeb AI GatewaySWT-APIZetaTechs API华际 APITommyLam API毫秒APIMentoe APIDibin84 API HubQQ Code123NHH APISeamee APIRinkoAISoul 公益站逆龙傲公益站wuer的api站MineWuer APIMapleLeaf APIAI APIDNSHEWAADRIMyNav AI | 165.73 t/s | 6.34 s | 15 |
ImageChatAsyncA high-resolution image generation model by Black Forest Labs in the FLUX.2 series, offering state-of-the-art quality and detail. | N/A | N/A | 0 |
128KTools33KChatZhipu GLM-4 is a flagship bilingual model from Zhipu AI with strong Chinese-English reasoning, tool use, and long-context chat capabilities. | N/A | N/A | 0 |
Intern-S1Free Intern-S1 is a language model in the InternLM series, offering general-purpose reasoning, code generation, and multimodal capabilities. | N/A | N/A | 0 |
ImageAlibaba Qwen Image is an image generation model in the Qwen series, offering advanced image generation capabilities. | N/A | N/A | 0 |
DeepSeek Coder is a code-specialized language model in the DeepSeek series, optimized for code generation, completion, and software development tasks. | N/A | N/A | 0 |
Alibaba Qwen3-Next is an ultra-efficient next-gen architecture (80B total, 3B active) with hybrid attention and sparse MoE, designed for long-context training and inference at dramatically lower cost. | N/A | N/A | 0 |
Open Weights32.8K32K向量模型Alibaba Qwen3 Embedding is an embedding model in the Qwen series, optimized for text embedding and similarity tasks. | N/A | N/A | 0 |
ReasoningToolsOpen WeightsVisionAlibaba Qwen3 VL is a vision-language model in the Qwen series, offering advanced multimodal capabilities for image and text understanding. | N/A | N/A | 0 |
128KReasoningToolsOpen WeightsDeepSeek V3 is DeepSeek flagship MoE language model with 671B total parameters, delivering strong performance in reasoning, coding, and multilingual tasks at competitive inference cost. | 26.71 t/s | 3.41 s | 5 |
CodexReasoningToolsFilesOpenAI GPT-5.2 is a GPT-5 series model emphasizing advanced reasoning, multimodal understanding, and high-quality outputs for complex enterprise workloads. | 兔子API初叶🍂Furry API180txt API +50 moreShow fewerFuture Hubllm-2-apiKternaDMXAPIPrivnodeVenlacyLiuwang API933999 APIFeng Love APILLM API天宫造物Codex Easy丸美小沐写作丸美小沐MagicAIMitchll-APIAIO通用智能服务平台ZenScale AI汪汪中转站Liunew APIAIsa42公益站Codex APIDeepRouter随时跑路公益站Neb 公益站Koyeb AI GatewaySWT-APIZetaTechs APISmart APIKJK APITommyLam API毫秒APIDibin84 API HubSeamee APIRinkoAICHB APIJuCodeSoul 公益站逆龙傲公益站不知道叫啥Mars HKWzjself APIwzjself中转站AI APIDNSHEWAADRIMyNav AI至强API10dian-API | 64.56 t/s | 4.77 s | 5 |
ReasoningToolsFilesVisionOpenAI GPT-5.2 Pro is a high-capability language model in the GPT-5 series, offering enhanced reasoning, code generation, and multimodal capabilities. | N/A | N/A | 0 |
ReasoningToolsFilesOpen WeightsZhipu AI GLM-4.6V is a multimodal vision-language model in the GLM series, supporting both text and image understanding. | N/A | N/A | 0 |
CodexReasoningToolsFilesOpenAI GPT-5.1 Codex Max is a code-specialized variant in the GPT-5 series, optimized for code generation, debugging, and software development tasks. | N/A | N/A | 0 |
ReasoningTools163.8K164KDeepSeek V3.2 is an upgraded V3-series MoE model with stronger reasoning, coding, and math performance, widely available through OpenAI-compatible API relays. | 兔子APIZero APIFuture Hub +35 moreShow fewer霸气公益平台DMXAPIPrivnodeAI Claw APIEnenCloud API933999 APILLM API17NAS API情酱的API站CxyKevin APIMitchll-API百万APIAIO通用智能服务平台AIsaDeepRouterNeb 公益站Koyeb AI GatewaySWT-APIZetaTechs API毫秒APISeamee APIRinkoAICHB APIwuer的api站APDSMMineWuer API星见雅 APIMapleLeaf APIAI APIWAADRIC85 API温云10dian-APIGOU APIOpenOpen8 API | 36.08 t/s | 1.40 s | 5 |
ReasoningToolsFilesVisionAnthropic Claude Opus 4.5 is the flagship Claude model for the hardest reasoning, research, and multi-step agent tasks where accuracy matters more than speed. | 45.89 t/s | 2.49 s | 5 |
Olmo 3 32B Think is a large-scale, 32-billion-parameter model purpose-built for deep reasoning, complex logic chains and advanced instruction-following scenarios. Its capacity enables strong performance on demanding evaluation tasks and... | N/A | N/A | 0 |
EmbeddingAn English embedding model by BAAI, designed for generating text embeddings for retrieval and similarity tasks. | N/A | N/A | 0 |
CodexReasoningToolsFilesOpenAI GPT-5.1 is a GPT-5 series model focused on improved reasoning depth, tool use, and reliable long-form generation for production AI applications. | N/A | N/A | 0 |
CodexReasoningToolsFilesOpenAI GPT-5.1 Codex is a code-specialized variant in the GPT-5 series, optimized for code generation, debugging, and software development tasks. | N/A | N/A | 0 |
CodexReasoningToolsFilesOpenAI GPT-5.1 Codex Mini is a code-specialized variant in the GPT-5 series, optimized for code generation, debugging, and software development tasks. | N/A | N/A | 0 |
ReasoningTools262KOpen WeightsMoonshot AI Kimi K2 Thinking is a reasoning model in the Kimi series, designed for complex reasoning, problem-solving, and analytical tasks. | N/A | N/A | 0 |
Amazon Nova Premier V1 is a multimodal vision-language model in the Nova family, designed for high-quality image understanding, document analysis, and visual reasoning across production workloads. | N/A | N/A | 0 |
2KGoogle Gemini Embedding 001 is an embedding model, designed for generating vector representations of text for retrieval and semantic search. | N/A | N/A | 0 |
A model–provider pair whose current input and output prices are both $0 per token. Some providers offer permanent free tiers, others only give one-time credits to new accounts. We mark an offering as free only while its public pricing is zero, and re-check pricing pages regularly.
Most have rate limits — requests per minute, daily caps, or context-length limits — and many require account signup with a verified phone or payment method. Some are time-limited promotions. Always read the provider's terms and quotas before depending on a free endpoint.
Some entries are community-run relays ("公益站") that bundle paid upstream keys and redistribute access for free at the operator's expense. They often advertise larger quotas and a broader model list than official free tiers, but reliability is much lower: operators can pull the plug or disappear overnight, pricing and quotas can change without notice, and many sites are invite-only — requiring a GitHub invite, a forum referral, or a closed community to register. Some keep signups disabled indefinitely. Treat them as best-effort backup channels; keep anything important on official paid endpoints.
Speed varies by model and provider. Sort the table by Most tested for the most reliable benchmarks, or pick a model family from the chips above to drill in. Each row shows median tokens per second and first-token latency from real API tests.
We send identical prompts to each provider through a five-round stress test, count output tokens with tiktoken, and measure both throughput (tokens per second) and time to first token. Numbers are aggregated as medians to resist outliers and refresh on a regular cadence.
For prototypes, side projects, and low-traffic tools, yes. Production traffic will usually hit a rate limit quickly. Treat the free tier as an evaluation channel: validate the model and provider, then move to a paid endpoint with the same model when you scale.
Either no provider currently offers it for free, the free promotion ended, or it has not been benchmarked yet. Open the model's main page to compare paid options, or let us know about a missing free provider via the feedback link in the footer.