Sponsored byFusecodeEnterprise coding API for Claude Code, Codex, and model workflows.
LogoLMSpeed
  • Free
  • Models
  • Providers
  • Leaderboard
  • Docs
LogoLMSpeed
  1. Home
  2. Free Models
LogoLMSpeed

The best API speed test tool

GitHubGitHubTwitterX (Twitter)Email
Product
  • Features
  • Pricing
  • FAQ
Leaderboard
  • Overview
  • Speed Ranking
  • Latency Ranking
  • Health Ranking
  • Model Pricing
  • Model Speed
  • Reasoning
  • Coding
Models
  • All Models
  • GPT
  • Claude
  • Gemini
  • DeepSeek
  • Llama
  • Qwen
Free Models
  • All Free Models
  • Free GPT
  • Free Claude
  • Free Gemini
  • Free DeepSeek
  • Free Llama
  • Free Qwen
Tools
  • Speed Test
  • Provider Audit
Company
  • About
Resources
  • Provider Directory
  • Documentation
  • Public API
  • Botab
  • VidBee
Legal
  • Cookie Policy
  • Privacy Policy
  • Terms of Service
© 2026 LMSpeed All Rights Reserved.Made by Nexmoe with ❤️
Free modelsFree API keysFree redeem codesPublic-benefit providersMoonshotAIKimiDeepSeekDeepSeekClaudeClaudeGeminiGeminiOpenAIGPTMetaAILlamaQwenQwenMistralMistralChatGLMGLMAzurePhiGemmaGemma

Free LLM API Models

Discover free LLM API models across providers with real speed and latency benchmarks. · Free models 2768 · Providers 155 · Free offerings 8686

2461–2520 of 2768
ModelProvidersSpeedLatencyTests
QwenAPEX-qwen-3.5-plus阿里巴巴Free
N/AN/A0
MoonshotAIAPEX-kimi-k2-thinkingMoonshotFree
N/AN/A0
MoonshotAIAPEX-kimi-k2.5MoonshotFree
N/AN/A0

LMSpeed tracks 2768 LLM models available for free across 155 API providers. Free tiers vary by provider — some offer limited daily requests, others provide free credits for new users. All speed data is from real API tests.

  • 1
  • 41
  • 42
  • 43
  • 47

How to use

Find a free model, compare its providers, then open the provider or model detail page before you start testing.

  1. 1

    Filter by model family

    Use search, model-family filters, and capability tags to narrow the directory to the free models you actually want to try.

  2. 2

    Compare speed and latency

    Check provider count, benchmark volume, throughput, and first-token latency so you are not choosing from price alone.

  3. 3

    Open the provider details

    Review the provider page for availability, health checks, pricing notes, and any free-tier limits before building on it.

Free LLM API FAQ

Common questions about free model tiers and how to use this directory.

MinimaxAPEX-minimax-m2MinimaxFree
10dian-API
N/A
N/A
0
MinimaxAPEX-minimax-m2.5MinimaxFree
10dian-API
N/AN/A0
TTSOpenAIFree
10dian-API
N/AN/A0
Qwenqwen3 Vision Model阿里巴巴Free
10dian-API
N/AN/A0
MoonshotAIkimi-k2-thinking-agent(次模型)MoonshotFree
10dian-API
N/AN/A0
GeminiAPEX-gemini-2.5-flash-liteGoogleFree
10dian-API
N/AN/A0
GeminiB-gemini-3.1-pro-preview-thinkingGoogleFree
10dian-API
N/AN/A0
GeminiB-gemini-3.1-pro-preview-maxthinkingGoogleFree
10dian-API
N/AN/A0
GeminiB-gemini-3.1-pro-previewGoogleFree
10dian-API
N/AN/A0
DeepSeekAPEX-deepseek-v3.2-chatDeepSeekFree
10dian-API
N/AN/A0
DeepSeekAPEX-deepseek-r1DeepSeekFree
10dian-API
N/AN/A0
coder-model-KC阿里巴巴Free
10dian-API
N/AN/A0
GeminiB-gemini-3-pro-preview-thinkingGoogleFree
10dian-API
N/AN/A0
GeminiB-gemini-3-pro-previewGoogleFree
10dian-API
N/AN/A0
GeminiB-gemini-3-flash-preview-thinkingGoogleFree
10dian-API
N/AN/A0
ClaudeAPEX-claude-sonnet-4-thinkingAnthropicFree
10dian-API
N/AN/A0
DeepSeekAPEX-deepseek-v3DeepSeekFree
10dian-API
N/AN/A0
GeminiAPEX-gemini-2.5-flashGoogleFree
10dian-API
N/AN/A0
DeepSeekAPEX-deepseek-v3.1DeepSeekFree
10dian-API
N/AN/A0
DeepSeekAPEX-deepseek-v3.2DeepSeekFree
10dian-API
N/AN/A0
GeminiB-gemini-3-flash-previewGoogleFree
10dian-API
N/AN/A0
GeminiB-gemini-2.5-pro-thinkingGoogleFree
10dian-API
N/AN/A0
GeminiB-gemini-2.5-proGoogleFree
10dian-API
N/AN/A0
GeminiB-gemini-2.5-flashGoogleFree
10dian-API
N/AN/A0
QwenAPEX-qwen3-vl-plus阿里巴巴Free
10dian-API
N/AN/A0
MiMo-V2-TTSXiaomi Token Plan (Singapore)Free
Open WeightsAudio8.2K8KXiaomi MiMo-V2-TTS is a text-to-speech model in the MiMo series, optimized for natural speech synthesis and voice generation tasks.
Dext APIDMXAPI17NAS API
+5 moreShow fewer情酱的API站42公益站Smart APITommyLam APIWAADRI
N/AN/A0
ClaudeClaude Opus 4.7 MaxAnthropicFree
Anthropic Claude Opus 4.7 Max is a high-capability language model in the Claude series, offering enhanced reasoning, code generation, and multimodal capabilities.
Hank Workspace APIMyNav AI
N/AN/A0
Gemini[L]gemini-3-pro-previewGoogleFree
GOU API
N/AN/A0
Gemini[L]gemini-3.1-flash-image-previewGoogleFree
GOU API
N/AN/A0
Gemini[L]gemini-3-flash-previewGoogleFree
GOU API
N/AN/A0
DeepSeekDeepSeek V4DeepSeekFree
DeepSeek V4 is DeepSeek's next-generation foundation model family, built for advanced reasoning, coding, math, and long-context agent applications.
Seamee API
N/AN/A0
QwenQwen1.8B Long ContextAlibabaFree
Alibaba Qwen1.8B Long Context is a compact language model in the Qwen series, optimized for low-latency responses and efficient inference.
DMXAPI毫秒API
N/AN/A0
QwenQwen1.8BAlibabaFree
Alibaba Qwen1.8B is a compact language model in the Qwen series, optimized for low-latency responses and efficient inference.
DMXAPI毫秒API
N/AN/A0
MetaAILlama 3.1MetaFree
Meta Llama 3.1 extends the Llama 3 family with stronger reasoning, tool use, and long-context support across 8B to 405B scales.
兔子APILLM APISmart API
+1 moreShow fewer毫秒API
N/AN/A0
MetaAILlama 3.3MetaFree
ToolsOpen Weights128KMeta Llama 3.3 is an updated Llama 3 open model with improved instruction following, multilingual support, and efficient inference.
Mitchll-APICHB API
N/AN/A0
MetaAIDracarys Llama 3.1 InstructMetaFree
Dracarys Llama 3.1 Instruct is an instruction-tuned variant, optimized for following instructions and conversational tasks.
Dext API初叶🍂Furry API91VIP API
+5 moreShow fewerAI Claw APISWT-API星见雅 APIC85 API温云
N/AN/A0
MetaAIUsdcode Llama 3.1 InstructMetaFree
Usdcode Llama 3.1 Instruct is an instruction-tuned variant, optimized for following instructions and conversational tasks.
AI Claw APIMitchll-API星见雅 API
N/AN/A0
QwenDeepSeek R1 Distill Qwen 1.5BDeepSeekFree
DeepSeek R1 Distill Qwen 1.5B is a reasoning model in the DeepSeek series, designed for complex reasoning, problem-solving, and analytical tasks.
Seamee API10dian-API
N/AN/A0
MetaAIMeta Llama 3.1 InstructMetaFree
Tools33KOpen Weights128KMeta Llama 3.1 Instruct is an instruction-tuned variant in the Llama series, optimized for following instructions and conversational tasks.
LLM APISWT-APICHB API
N/AN/A0
DeepSeekDeepSeek V2.5DeepSeekFree
DeepSeek V2.5 is a language model in the DeepSeek series, offering general-purpose reasoning, code generation, and multimodal capabilities.
Xiao WanWAADRI
N/AN/A0
DeepSeekDeepSeek V3.1DeepSeekFree
ReasoningTools128KChatDeepSeek V3.1 is an open-weights style frontier model from DeepSeek with strong math, coding, and Chinese-English bilingual reasoning.
兔子APIZero API霸气公益平台
+18 moreShow fewerDMXAPIPrivnode933999 APILLM APIMitchll-API百万APIAIO通用智能服务平台AIsaDeepRouterZetaTechs API毫秒APIwuer的api站APDSMMineWuer API星见雅 APIMapleLeaf APIWAADRI10dian-API
N/AN/A0
MinimaxMiniMax FileMinimaxFree
MiniMax File is the MiniMax API endpoint for retrieving and managing uploaded files and generated outputs such as video clips and synthesized speech assets.
神马中转API小豆包APINewagiai
+2 moreShow fewer柏拉图AI毫秒API
N/AN/A0
QwenQwen3.5 Plus Image Edit阿里巴巴Free
Alibaba Qwen3.5 Plus Image Edit is an image editing model in the Qwen 3.5 series, supporting instruction-based image modification, inpainting, and visual refinement tasks.
MapleLeaf APIDNSHE
N/AN/A0
GeminiGemini 2.5 Pro DeepSearchGoogleFree
Google Gemini 2.5 Pro DeepSearch is a search-augmented language model in the Gemini series, integrating web retrieval to provide up-to-date answers.
兔子API
N/AN/A0
DeepSeekDeepSeek V3 TurboDeepSeekFree
DeepSeek V3 Turbo is a throughput-optimized DeepSeek V3 tier for large-scale chat, coding assistance, and MoE inference with strong price-to-performance.
AIO通用智能服务平台
N/AN/A0
MinimaxMiniMax Voice DesignMinimaxFree
MiniMax Voice Design is a voice synthesis model that creates custom voices from stylistic prompts, enabling personalized TTS profiles for chatbots, narration, and interactive applications.
毫秒API
N/AN/A0
OpenAIGPT-4o StudyOpenAIFree
ChatOpenAI GPT-4o Study is an instruction-tuned variant in the GPT-4 series, optimized for following instructions and conversational tasks.
兔子API毫秒API
N/AN/A0
GeminiGemini 1.5 Pro 002GoogleFree
Google Gemini 1.5 Pro 002 is a high-capability language model in the Gemini series, offering enhanced reasoning, code generation, and multimodal capabilities.
兔子API
N/AN/A0
DeepSeekDeepSeek R1 TurboDeepSeekFree
DeepSeek R1 Turbo is a reasoning model in the DeepSeek series, designed for complex reasoning, problem-solving, and analytical tasks.
AIO通用智能服务平台
N/AN/A0
GeminiGemini 2.5 Flash DeepSearchGoogleFree
Google Gemini 2.5 Flash DeepSearch is a search-augmented language model in the Gemini series, integrating web retrieval to provide up-to-date answers.
兔子API
N/AN/A0
MinimaxMiniMax File UploadMinimaxFree
MiniMax File Upload is the MiniMax API file-management endpoint for uploading audio, video, and image assets used in voice cloning, video generation, and multimodal understanding workflows.
毫秒API
N/AN/A0
ClaudeClaude 2.0AnthropicFree
文本Anthropic Claude 2.0 introduced stronger reasoning and safer long-form generation for early production Claude deployments.
兔子APIMagicAI
N/AN/A0
QwenQwen3.5 Omni Flash阿里巴巴Free
Alibaba Qwen3.5 Omni Flash is a multimodal model in the Qwen series, offering text, image, video, and audio understanding capabilities.
DMXAPI42公益站Koyeb AI Gateway
+3 moreShow fewerSWT-APIMapleLeaf APIWAADRI
N/AN/A0
ChatGLMGLM-4 FlashXZhipu AIFree
Zhipu AI GLM-4 FlashX is an accelerated GLM-4 Flash build for sub-second responses in customer support bots and high-QPS API gateways.
MapleLeaf API
N/AN/A0
GeminiGemini Robotics ER 1.5GoogleFree
Google Gemini Robotics ER 1.5 is a robotics-focused multimodal model in the Gemini series, designed for embodied reasoning and physical task planning.
Koyeb AI GatewaySWT-API
N/AN/A0
MinimaxMiniMax Voice CloneMinimaxFree
MiniMax Voice Clone is a voice cloning model that replicates speaker characteristics from short reference audio, enabling personalized TTS for dubbing, assistants, and creative media production.
毫秒API
N/AN/A0
QwenQwen3.5 Omni Plus阿里巴巴Free
Alibaba Qwen3.5 Omni Plus is the flagship omnimodal model in the Qwen series, natively understanding text, images, audio, and video with a 256K context window, real-time speech interaction, and agentic tool use.
DMXAPI42公益站Koyeb AI Gateway
+3 moreShow fewerSWT-APIMapleLeaf APIWAADRI
N/AN/A0
What counts as a free LLM API on LMSpeed?

A model–provider pair whose current input and output prices are both $0 per token. Some providers offer permanent free tiers, others only give one-time credits to new accounts. We mark an offering as free only while its public pricing is zero, and re-check pricing pages regularly.

Are these free APIs really free? Any catch?

Most have rate limits — requests per minute, daily caps, or context-length limits — and many require account signup with a verified phone or payment method. Some are time-limited promotions. Always read the provider's terms and quotas before depending on a free endpoint.

Are community-run relays and non-profit aggregators included? Any extra caveats?

Some entries are community-run relays ("公益站") that bundle paid upstream keys and redistribute access for free at the operator's expense. They often advertise larger quotas and a broader model list than official free tiers, but reliability is much lower: operators can pull the plug or disappear overnight, pricing and quotas can change without notice, and many sites are invite-only — requiring a GitHub invite, a forum referral, or a closed community to register. Some keep signups disabled indefinitely. Treat them as best-effort backup channels; keep anything important on official paid endpoints.

Which free LLM API is fastest?

Speed varies by model and provider. Sort the table by Most tested for the most reliable benchmarks, or pick a model family from the chips above to drill in. Each row shows median tokens per second and first-token latency from real API tests.

How does LMSpeed measure speed and latency?

We send identical prompts to each provider through a five-round stress test, count output tokens with tiktoken, and measure both throughput (tokens per second) and time to first token. Numbers are aggregated as medians to resist outliers and refresh on a regular cadence.

Can I use a free LLM API in production?

For prototypes, side projects, and low-traffic tools, yes. Production traffic will usually hit a rate limit quickly. Treat the free tier as an evaluation channel: validate the model and provider, then move to a paid endpoint with the same model when you scale.

Why don't I see a specific model in this list?

Either no provider currently offers it for free, the free promotion ended, or it has not been benchmarked yet. Open the model's main page to compare paid options, or let us know about a missing free provider via the feedback link in the footer.

10dian-API
10dian-API
10dian-API
LMSpeed community

Join the LMSpeed QQ Channel

Get free-model updates, share API keys and redeem codes, and chat with other LMSpeed users.

Channel IDpd73677591
QR code for joining the LMSpeed QQ Channel
Scan with QQ to join

Latest LLM models

Explore the latest published LLMs tracked by LMSpeed, then open a model to compare providers, pricing, benchmarks, and free availability.

Browse all models
  • OpenAIGPT-5.6 LunaOpenAI

    GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

    Released Jul 9, 2026106 providers1 free
  • OpenAIGPT-5.6 TerraOpenAI

    GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

    Released Jul 9, 2026114 providers1 free
  • OpenAIGPT-5.6 SolOpenAI

    GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

    Released Jul 9, 2026116 providers1 free
  • GrokGrok 4.5xAI

    Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

    Released Jul 8, 202670 providers2 free
  • HunyuanHy3Tencent

    Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...

    Released Jul 6, 202618 providers2 free
  • ClaudeClaude Sonnet 5Anthropic

    Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

    Released Jun 30, 2026105 providers1 free