Sponsored byFusecodeEnterprise coding API for Claude Code, Codex, and model workflows.
LogoLMSpeed
  • Free
  • Models
  • Providers
  • Leaderboard
  • Docs
LogoLMSpeed
  1. Home
  2. Free Models
LogoLMSpeed

The best API speed test tool

GitHubGitHubTwitterX (Twitter)Email
Product
  • Features
  • Pricing
  • FAQ
Leaderboard
  • Overview
  • Speed Ranking
  • Latency Ranking
  • Health Ranking
  • Model Pricing
  • Model Speed
  • Reasoning
  • Coding
Models
  • All Models
  • GPT
  • Claude
  • Gemini
  • DeepSeek
  • Llama
  • Qwen
Free Models
  • All Free Models
  • Free GPT
  • Free Claude
  • Free Gemini
  • Free DeepSeek
  • Free Llama
  • Free Qwen
Tools
  • Speed Test
  • Provider Audit
Company
  • About
Resources
  • Provider Directory
  • Documentation
  • Public API
  • Botab
  • VidBee
Legal
  • Cookie Policy
  • Privacy Policy
  • Terms of Service
© 2026 LMSpeed All Rights Reserved.Made by Nexmoe with ❤️
Free modelsFree API keysFree redeem codesPublic-benefit providersMoonshotAIKimiDeepSeekDeepSeekClaudeClaudeGeminiGeminiOpenAIGPTMetaAILlamaQwenQwenMistralMistralChatGLMGLMAzurePhiGemmaGemma

Free LLM API Models

Discover free LLM API models across providers with real speed and latency benchmarks. · Free models 2773 · Providers 157 · Free offerings 8727

2521–2580 of 2773
ModelProvidersSpeedLatencyTests
Qwen阿里巴巴Free
Alibaba Qwen3.5 Omni Flash is a multimodal model in the Qwen series, offering text, image, video, and audio understanding capabilities.
+3 moreShow fewer
N/AN/A0
ChatGLMZhipu AIFree
Zhipu AI GLM-4 FlashX is an accelerated GLM-4 Flash build for sub-second responses in customer support bots and high-QPS API gateways.
N/AN/A0
GeminiGoogle

LMSpeed tracks 2773 LLM models available for free across 157 API providers. Free tiers vary by provider — some offer limited daily requests, others provide free credits for new users. All speed data is from real API tests.

  • 1
  • 42
  • 43
  • 44
  • 47

How to use

Find a free model, compare its providers, then open the provider or model detail page before you start testing.

  1. 1

    Filter by model family

    Use search, model-family filters, and capability tags to narrow the directory to the free models you actually want to try.

  2. 2

    Compare speed and latency

    Check provider count, benchmark volume, throughput, and first-token latency so you are not choosing from price alone.

  3. 3

    Open the provider details

    Review the provider page for availability, health checks, pricing notes, and any free-tier limits before building on it.

Free LLM API FAQ

Common questions about free model tiers and how to use this directory.

Free
Google Gemini Robotics ER 1.5 is a robotics-focused multimodal model in the Gemini series, designed for embodied reasoning and physical task planning.
Koyeb AI GatewaySWT-API
N/A
N/A
0
MinimaxMiniMax Voice CloneMinimaxFree
MiniMax Voice Clone is a voice cloning model that replicates speaker characteristics from short reference audio, enabling personalized TTS for dubbing, assistants, and creative media production.
SmokeDivine AI简易-API中转站N1N
+5 moreShow fewer向量引擎Jeniya AI APIZhongzhuan ChatYUNWU API毫秒API
N/AN/A0
QwenQwen3.5 Omni Plus阿里巴巴Free
Alibaba Qwen3.5 Omni Plus is the flagship omnimodal model in the Qwen series, natively understanding text, images, audio, and video with a 256K context window, real-time speech interaction, and agentic tool use.
DMXAPI42公益站Koyeb AI Gateway
+3 moreShow fewerSWT-APIMapleLeaf APIWAADRI
N/AN/A0
DeepSeekDeepSeek V3.2 ReasonerDeepSeekFree
DeepSeek V3.2 Reasoner is a reasoning model in the DeepSeek series, designed for complex reasoning, problem-solving, and analytical tasks.
wuer的api站MineWuer API
N/AN/A0
MinimaxMiniMax M2.1 LightningFree
MiniMax M2.1 Lightning is an ultra-fast MiniMax tier for real-time voice companions, streaming chat, and latency-sensitive multimodal experiences.
DMXAPI42公益站
N/AN/A0
GeminiGemini 2.5 Flash LiveGoogleFree
Google Gemini 2.5 Flash Live is a realtime audio model in the Gemini series, supporting low-latency speech and conversational interactions.
WAADRI
N/AN/A0
OpenAIGPT-4o DALL-EOpenAIFree
OpenAI GPT-4o DALL-E is an image generation model combining GPT-4o reasoning with DALL-E image synthesis, enabling detailed visual creation from conversational prompts.
兔子API
N/AN/A0
ChatGLMGLM-4.1v Thinking FlashZhipu AIFree
Zhipu AI GLM-4.1v Thinking Flash is a reasoning model in the GLM series, designed for complex reasoning, problem-solving, and analytical tasks.
Koyeb AI GatewaySWT-API毫秒API
+1 moreShow fewerSeamee API
N/AN/A0
GeminiGemini 2.5 Pro 1MGoogleFree
Google Gemini 2.5 Pro 1M is a high-capability language model in the Gemini series with a 1 million-token context window, offering enhanced reasoning, code generation, and multimodal capabilities.
Xiao Wan
N/AN/A0
QwenQwen3.6 Plus Thinking阿里巴巴Free
Alibaba Qwen3.6 Plus Thinking is a reasoning-focused variant in the Qwen series, designed for complex reasoning and problem-solving tasks.
Mitchll-API42公益站Koyeb AI Gateway
+3 moreShow fewerSWT-APISmart APIMapleLeaf API
N/AN/A0
GeminiGemini 3.1 FlashGoogleFree
Google Gemini 3.1 Flash extends the Gemini 3 Flash line with improved reasoning and multimodal accuracy for production assistants and search-augmented apps.
丸美小沐写作丸美小沐AIO通用智能服务平台
+1 moreShow fewerWAADRI
N/AN/A0
ChatGLMGLM-4 AirXZhipu AIFree
Zhipu AI GLM-4 AirX is an enhanced lightweight GLM-4 Air variant with higher throughput for enterprise chatbots and bilingual knowledge bases.
兔子APIDMXAPI
N/AN/A0
GeminiGemini EmbeddingGoogleFree
Google Gemini Embedding is an embedding model, designed for generating vector representations of text for retrieval and semantic search.
兔子API
N/AN/A0
GeminiGemini 2.0 Flash Live 001GoogleFree
Google Gemini 2.0 Flash Live 001 is a realtime audio model in the Gemini series, supporting low-latency speech and conversational interactions.
兔子API
N/AN/A0
MiMo-V2-OmniMiMoFree
ReasoningToolsFilesOpen WeightsXiaomi MiMo-V2-Omni is the omnimodal model in the V2 series on the Xiaomi MiMo API platform, supporting text, image, video, and audio understanding within a unified architecture. Pricing: 1x token consumption (baseline).
Dext APIZero APIDMXAPI
+10 moreShow fewer933999 API17NAS API情酱的API站42公益站Smart APITommyLam API毫秒APIAI APIDNSHEWAADRI
N/AN/A0
OpenAIGPT-3.5 NetOpenAIFree
OpenAI GPT-3.5 Net is a language model in the GPT-3.5 series, offering general-purpose reasoning, code generation, and multimodal capabilities.
毫秒API
N/AN/A0
MinimaxMiniMax M2.5 HighSpeedMiniMax Coding Plan (minimax.io)Free
ReasoningToolsOpen Weights204.8KMiniMax M2.5 HighSpeed prioritizes token throughput and low latency for large-scale conversational AI, content generation, and API relay traffic.
Zero APIOptAITommyLam API
+2 moreShow fewer毫秒APIMars HK
N/AN/A0
ChatGLMGLM-4V FlashZhipu AIFree
Zhipu AI GLM-4V Flash is a multimodal vision-language model in the GLM series, supporting both text and image understanding.
ocool AIMitchll-API百万API
+2 moreShow fewerKoyeb AI GatewaySWT-API
N/AN/A0
GeminiGemini Live 2.5 FlashGoogleFree
Google Gemini Live 2.5 Flash is a realtime audio model in the Gemini series, supporting low-latency speech and conversational interactions.
兔子API
N/AN/A0
Arctic Embed LFree
Arctic Embed L is an embedding model, designed for generating vector representations of text for retrieval and semantic search.
Dext API初叶🍂Furry APIAI Claw API
+5 moreShow fewerKoyeb AI GatewaySWT-API星见雅 APIC85 API温云
N/AN/A0
ChatGLMGLM-4.1v Thinking FlashXZhipu AIFree
FilesVision64KZhipu AI GLM-4.1v Thinking FlashX is a reasoning model in the GLM series, designed for complex reasoning, problem-solving, and analytical tasks.
毫秒API不知道叫啥MapleLeaf API
N/AN/A0
QwenQwen3.5 Max阿里巴巴Free
Alibaba Qwen3.5 Max is a high-capability language model in the Qwen series, offering enhanced reasoning, code generation, and multimodal capabilities.
42公益站MapleLeaf API
N/AN/A0
QwenQwen3.5 Plus Thinking通义千问Free
Alibaba Qwen3.5 Plus Thinking is a reasoning-focused variant in the Qwen series, designed for complex reasoning and problem-solving tasks.
42公益站Koyeb AI GatewaySWT-API
+4 moreShow fewerZetaTechs APISeamee APIMapleLeaf APIDNSHE
N/AN/A0
AzurePhi 3.5 MoE InstructGitHub ModelsFree
ReasoningToolsOpen Weights128KMicrosoft Phi 3.5 MoE Instruct is a mixture-of-experts instruction-tuned variant in the Phi series, optimized for following instructions and conversational tasks.
Dext APIChooseC API初叶🍂Furry API
+8 moreShow fewerWSocket AIChooseC APIAI Claw APIKoyeb AI GatewaySWT-API星见雅 APIC85 API温云
N/AN/A0
ChatGLMGLM-4.5-XZhipu AIFree
Zhipu GLM-4.5-X is a premium GLM-4.5 variant optimized for complex agentic reasoning, tool orchestration, and higher-quality instruction following.
兔子APIMitchll-API毫秒API
N/AN/A0
ClaudeClaude Opus 4.6 MaxAnthropicFree
Anthropic Claude Opus 4.6 Max is a high-capability language model in the Claude series, offering enhanced reasoning, code generation, and multimodal capabilities.
ZetaTechs APIMyNav AI
N/AN/A0
Italia InstructFree
Italia Instruct is an instruction-tuned language model, optimized for following instructions and conversational tasks.
APDSM
N/AN/A0
MinimaxMiniMax M2.1 HighSpeedMinimaxFree
ChatMiniMax M2.1 HighSpeed is a performance-tuned MiniMax M2.1 build for high-concurrency chat, roleplay, and streaming text generation.
TommyLam API毫秒APIMars HK
N/AN/A0
MoonshotAIKimi K2 TurboMoonshot AIFree
ToolsOpen Weights262.1KMoonshot AI Kimi K2 Turbo is a fast Kimi K2 tier for long-context chat, document analysis, and Chinese-first assistants with competitive inference speed.
ZetaTechs APIWAADRI
N/AN/A0
ChatGLMGLM-4 FlashZhipu AIFree
128KZhipu AI GLM-4 Flash is a budget-friendly GLM-4 variant for everyday dialogue, lightweight coding help, and high-volume API consumption.
兔子API6655 翻译小站Chlink API
+10 moreShow fewerMIX APIIXIOCCAPIDMXAPIMitchll-API百万APIKoyeb AI GatewaySWT-APISeamee API不知道叫啥MapleLeaf API
N/AN/A0
DeepSeekDeepSeek Coder InstructDeepSeekFree
DeepSeek Coder Instruct is a code-specialized variant in the DeepSeek series, optimized for code generation, debugging, and software development tasks.
91VIP APIReal AI WANSWT-API
+2 moreShow fewerC85 API温云
N/AN/A0
ChatGLMGLM-4 AirZhipu AIFree
128KZhipu AI GLM-4 Air is a compact GLM-4 model for mobile assistants, on-device gateways, and cost-sensitive bilingual chat applications.
兔子APIDMXAPI不知道叫啥
+1 moreShow fewerMapleLeaf API
N/AN/A0
Solar InstructFree
Upstage Solar Instruct is an instruction-tuned language model in the Solar series, optimized for chat, summarization, and enterprise document workflows with strong Korean and English performance.
SWT-APIAPDSMMapleLeaf API
+2 moreShow fewerC85 API温云
N/AN/A0
PanGu Pro MoESiliconFlow (China)Free
Reasoning128KReasoningHuawei PanGu Pro MoE is a mixture-of-experts language model designed for enterprise Chinese-English tasks, reasoning, and scalable cloud inference.
Xiao WanWAADRI
N/AN/A0
ChatGLMGLM-4 PlusZhipu AIFree
Zhipu GLM-4 Plus is a high-end GLM-4 tier optimized for complex dialogue, analysis, and enterprise assistant scenarios.
DMXAPISmart API毫秒API
N/AN/A0
ChatGLMGLM-4.6V FlashZhipu AIFree
ReasoningToolsFilesOpen WeightsZhipu AI GLM-4.6V Flash is a multimodal vision-language model in the GLM series, supporting both text and image understanding.
Dext APIWSocket AIClaw API
+10 moreShow fewer小天公益站VSLLMMoyanjdc APIFuture HubKternaKoyeb AI GatewaySWT-APISeamee APIMapleLeaf APIGOU API
N/AN/A0
TeleSpeechASRFree
TeleSpeechASR is a speech-to-text model designed for accurate audio transcription and recognition, supporting voice input pipelines, meeting notes, and accessibility applications.
Koyeb AI GatewaySWT-APIWAADRI
N/AN/A0
MetaAILlama 3.2 Nemoretriever 300m Embed v1MetaFree
Meta Llama 3.2 Nemoretriever 300m Embed v1 is an embedding model, designed for generating vector representations of text for retrieval and semantic search.
AI Claw APIKoyeb AI GatewaySWT-API
+3 moreShow fewer星见雅 APIC85 API温云
N/AN/A0
AzurePhi 4 Multimodal InstructGitHub ModelsFree
128KReasoningToolsOpen WeightsMicrosoft Phi 4 Multimodal Instruct is a multimodal instruction-tuned variant in the Phi series, optimized for following instructions and conversational tasks.
Dext APIChooseC APIWSocket AI
+9 moreShow fewerChooseC APIAI Claw APIKoyeb AI GatewaySWT-API毫秒API星见雅 APIMapleLeaf APIC85 API温云
N/AN/A0
StepfunStep3SiliconFlow (China)Free
Tools文件Multimodal66KStepFun Step3 is a general-purpose language model with strong reasoning and instruction-following capabilities, designed for chat, coding assistance, and agentic workflows on the StepFun platform.
Xiao WanSeamee API
N/AN/A0
Jamba 1.5 Large InstructFree
AI21 Jamba 1.5 Large Instruct is a large instruction-tuned language model, optimized for following instructions and conversational tasks.
Dext API初叶🍂Furry APIAI Claw API
+5 moreShow fewerKoyeb AI GatewaySWT-API星见雅 APIC85 API温云
N/AN/A0
ChatGLMGLM-Z1 AirZhipu AIFree
Zhipu AI GLM-Z1 Air is a lightweight reasoning-oriented GLM-Z1 variant for math, logic puzzles, and structured problem solving at low cost.
毫秒API
N/AN/A0
ChatGLMGLM-4.5 FlashZhipu AIFree
ReasoningToolsOpen Weights131.1KZhipu AI GLM-4.5 Flash is a speed-first GLM-4.5 model for interactive chat, function calling, and bilingual customer-service automation.
兔子API6655 翻译小站Zero API
+10 moreShow fewer初叶🍂Furry API小天公益站Future HubKternaMitchll-APIKoyeb AI GatewaySWT-API毫秒APISeamee APIWAADRI
N/AN/A0
MinimaxMiniMax Hailuo 02MinimaxFree
VideoAsyncMiniMax Hailuo 02 is a video generation model that produces short video clips from text or image prompts, supporting creative content workflows on the MiniMax platform.
DMXAPI毫秒API
N/AN/A0
QwenQwen3.0AlibabaFree
Alibaba Qwen3.0 is an early Qwen3 generation model offering solid multilingual chat, instruction following, and cost-efficient API deployment.
933999 APIZetaTechs API毫秒API
+1 moreShow fewerSeamee API
N/AN/A0
Ring Flash 2.0SiliconFlow (China)Free
ReasoningTools131KReasoningRing Flash 2.0 is a fast and efficient language model, optimized for quick responses and high throughput.
Xiao WanWAADRI
N/AN/A0
FLUX DevFluxFree
Open WeightsVision77ImageAn open-weight image generation model by Black Forest Labs in the FLUX series, designed for development and experimentation.
毫秒APICHB API
N/AN/A0
ClaudeClaude 3 OpusAnthropicFree
ChatVisionAnthropic Claude 3 Opus was the top-tier Claude 3 model for complex analysis, long documents, and nuanced generation before later Opus revisions.
AIO通用智能服务平台Smart API毫秒API
N/AN/A0
GeminiGemini 1.5 FlashGoogleFree
ToolsFilesVisionAudioGoogle Gemini 1.5 Flash is a multimodal model with up to 1M tokens of context, ideal for long-document summarization, audio transcription, and vision tasks.
AIO通用智能服务平台CHB API
N/AN/A0
BGE Large ZH V1.5OtherFree
EmbeddingBGE Large ZH V1.5 is an embedding model, designed for generating vector representations of text for retrieval and semantic search.
6655 翻译小站初叶🍂Furry APIFuture Hub
+4 moreShow fewerKoyeb AI GatewaySWT-API毫秒APIWAADRI
N/AN/A0
GeminiGemini 2.0 ProGoogleFree
Google Gemini 2.0 Pro is a high-capability language model in the Gemini series, offering enhanced reasoning, code generation, and multimodal capabilities.
兔子APIAIO通用智能服务平台CHB API
N/AN/A0
MoonshotAIKimi K2 Thinking TurboMoonshot AIFree
ReasoningToolsOpen Weights262.1KMoonshot AI Kimi K2 Thinking Turbo is a reasoning model in the Kimi series, designed for complex reasoning, problem-solving, and analytical tasks.
DMXAPIZetaTechs APIWAADRI
N/AN/A0
GeminiGemini 2.0 Flash LiteGoogleFree
Google Gemini 2.0 Flash Lite is a lightweight and cost-efficient language model in the Gemini series, optimized for fast responses at reduced cost.
兔子APIKoyeb AI GatewaySWT-API
+2 moreShow fewer毫秒APICHB API
N/AN/A0
GeminiGemini 2.0 Flash Lite 001GoogleFree
Google Gemini 2.0 Flash Lite 001 is a compact language model in the Gemini series, optimized for low-latency responses and efficient inference.
兔子APIKoyeb AI GatewaySWT-API
+2 moreShow fewer毫秒APIAI API
N/AN/A0
ClaudeClaude 3 SonnetAnthropicFree
ChatVisionAnthropic Claude 3 Sonnet balanced capability and efficiency across text and vision tasks in the Claude 3 family.
AIO通用智能服务平台毫秒API
N/AN/A0
ClaudeClaude 3.7 SonnetAnthropicFree
ReasoningTools200KFilesAnthropic Claude 3.7 Sonnet improves reasoning depth and coding quality over earlier Sonnet releases while staying suitable for interactive products.
Mitchll-APIAIO通用智能服务平台Fangyuan API
+4 moreShow fewer毫秒APISeamee APICHB APIAI API
N/AN/A0
OpenAIGPT-4 DALL-EOpenAIFree
OpenAI GPT-4 DALL-E is an image generation model integrated with GPT-4 capabilities, producing detailed images from natural language descriptions for design and content workflows.
兔子API毫秒API
N/AN/A0
GeminiGemini 2.0 FlashGoogleFree
ToolsFilesVisionAudioGoogle Gemini 2.0 Flash introduces native tool use and agentic capabilities with fast multimodal inference for production chat and automation.
兔子API艾可APIAIO通用智能服务平台
+8 moreShow fewerDeepRouterKoyeb AI GatewaySWT-APISmart API毫秒APICHB APIMapleLeaf APIWAADRI
N/AN/A0
What counts as a free LLM API on LMSpeed?

A model–provider pair whose current input and output prices are both $0 per token. Some providers offer permanent free tiers, others only give one-time credits to new accounts. We mark an offering as free only while its public pricing is zero, and re-check pricing pages regularly.

Are these free APIs really free? Any catch?

Most have rate limits — requests per minute, daily caps, or context-length limits — and many require account signup with a verified phone or payment method. Some are time-limited promotions. Always read the provider's terms and quotas before depending on a free endpoint.

Are community-run relays and non-profit aggregators included? Any extra caveats?

Some entries are community-run relays ("公益站") that bundle paid upstream keys and redistribute access for free at the operator's expense. They often advertise larger quotas and a broader model list than official free tiers, but reliability is much lower: operators can pull the plug or disappear overnight, pricing and quotas can change without notice, and many sites are invite-only — requiring a GitHub invite, a forum referral, or a closed community to register. Some keep signups disabled indefinitely. Treat them as best-effort backup channels; keep anything important on official paid endpoints.

Which free LLM API is fastest?

Speed varies by model and provider. Sort the table by Most tested for the most reliable benchmarks, or pick a model family from the chips above to drill in. Each row shows median tokens per second and first-token latency from real API tests.

How does LMSpeed measure speed and latency?

We send identical prompts to each provider through a five-round stress test, count output tokens with tiktoken, and measure both throughput (tokens per second) and time to first token. Numbers are aggregated as medians to resist outliers and refresh on a regular cadence.

Can I use a free LLM API in production?

For prototypes, side projects, and low-traffic tools, yes. Production traffic will usually hit a rate limit quickly. Treat the free tier as an evaluation channel: validate the model and provider, then move to a paid endpoint with the same model when you scale.

Why don't I see a specific model in this list?

Either no provider currently offers it for free, the free promotion ended, or it has not been benchmarked yet. Open the model's main page to compare paid options, or let us know about a missing free provider via the feedback link in the footer.

Qwen3.5 Omni Flash
DMXAPI
42公益站
Koyeb AI Gateway
SWT-API
MapleLeaf API
WAADRI
GLM-4 FlashX
MapleLeaf API
Gemini Robotics ER 1.5
LMSpeed community

Join the LMSpeed QQ Channel

Get free-model updates, share API keys and redeem codes, and chat with other LMSpeed users.

Channel IDpd73677591
QR code for joining the LMSpeed QQ Channel
Scan with QQ to join