Discover free LLM API models across providers with real speed and latency benchmarks.Free models 2774Providers 157Free offerings 8768
| Model | Providers | Speed | Latency | Tests |
|---|---|---|---|---|
GoogleFree ToolsFilesVisionAudioGoogle Gemini 2.0 Flash introduces native tool use and agentic capabilities with fast multimodal inference for production chat and automation. | +9 moreShow fewer | N/A | N/A | 0 |
OpenAIFree OpenAI GPT-4o Audio is an audio-capable language model in the GPT-4 series, supporting voice input and output alongside text. |
LMSpeed tracks 2774 LLM models available for free across 157 API providers. Free tiers vary by provider — some offer limited daily requests, others provide free credits for new users. All speed data is from real API tests.
Find a free model, compare its providers, then open the provider or model detail page before you start testing.
Use search, model-family filters, and capability tags to narrow the directory to the free models you actually want to try.
Check provider count, benchmark volume, throughput, and first-token latency so you are not choosing from price alone.
Review the provider page for availability, health checks, pricing notes, and any free-tier limits before building on it.
Common questions about free model tiers and how to use this directory.
| N/A |
| N/A |
| 0 |
Zhipu AI GLM-4V Plus is a multimodal vision-language model in the GLM series, supporting both text and image understanding. | N/A | N/A | 0 |
FilesOpen WeightsVision8.2KDeepSeek OCR is an OCR model, specialized in extracting and understanding text from images and documents. | N/A | N/A | 0 |
OpenAI GPT-4o Mini Audio is an audio-capable language model in the GPT-4 series, supporting voice input and output alongside text. | N/A | N/A | 0 |
Tools131.1KChatxAI Grok-2 is a frontier multimodal language model in the Grok series, with strong chat, coding, and reasoning capabilities, real-time X platform integration, and a 131K-token context window. | N/A | N/A | 0 |
Alibaba Qwen3.1 is a Qwen3 series model with improved reasoning and coding over Qwen3.0, suitable for general chat and lightweight agent tasks. | N/A | N/A | 0 |
MiniMax Video is MiniMax's text-to-video and image-to-video generation API for creating cinematic clips from prompts, with asynchronous processing and director-style camera controls. | N/A | N/A | 0 |
Fish Speech 1.5Free Fish Speech 1.5 is a text-to-speech model by Fish Audio that generates natural, expressive speech from text, supporting multilingual voice output for content creation and voice assistants. | N/A | N/A | 0 |
ChatZhipu AI GLM-3 Turbo is a legacy high-speed GLM-3 tier still used in older integrations requiring stable Chinese dialogue and function calling. | N/A | N/A | 0 |
ReasoningTools128K131KMoonshot Kimi K2 Instruct is an instruction-tuned variant in the Kimi series, optimized for following instructions and conversational tasks. | 66.42 t/s | 0.82 s | 10 |
OpenAI GPT-4o Mini Realtime is a realtime audio model in the GPT-4 series, supporting low-latency speech and conversational interactions. | N/A | N/A | 0 |
ReasoningToolsOpen Weights204.8KMiniMax M2.7 HighSpeed is a throughput-optimized MiniMax M2.7 build for large-scale text and speech companion applications. | 54.34 t/s | 6.62 s | 5 |
131.1KToolsChatVisionxAI Grok 3 is a frontier language model in the Grok series with advanced reasoning, long-context understanding, and general-purpose assistant capabilities across text and multimodal tasks. | N/A | N/A | 0 |
Anthropic Claude 1.3 is a language model in the Claude series, offering general-purpose reasoning, code generation, and multimodal capabilities. | N/A | N/A | 0 |
Alibaba Qwen3 TTS Flash is a text-to-speech model, designed for high-quality speech synthesis from text input. | N/A | N/A | 0 |
ChatRealtime VoiceOpenAI GPT-4o Realtime is a realtime audio model in the GPT-4 series, supporting low-latency speech and conversational interactions. | N/A | N/A | 0 |
Google Gemini 2.0 Flash 001 is a pinned Gemini 2.0 Flash release for reproducible multimodal chat, coding, and agent pipelines. | N/A | N/A | 0 |
Open Weights512EmbeddingRerankBAAI BGE Reranker V2 M3 is a multilingual reranking model that reorders search and retrieval results by semantic relevance, widely used in RAG pipelines and hybrid search systems. | N/A | N/A | 0 |
Google Gemini 2.5 Flash Native Audio is an audio-capable language model in the Gemini series, supporting voice input and output alongside text. | N/A | N/A | 0 |
ChatVisionZhipu AI GLM-4V is a multimodal vision-language model in the GLM series, supporting both text and image understanding. | N/A | N/A | 0 |
Zhipu AI GLM-Z1 Flash is a fast reasoning model in the GLM-Z1 family for STEM tutoring, step-by-step logic, and lightweight agent loops. | N/A | N/A | 0 |
Zhipu AI GLM-Z1 AirX is an efficient GLM-Z1 Air variant tuned for high-QPS reasoning workloads and bilingual math-heavy chat. | N/A | N/A | 0 |
OpenAI GPT-4o Transcribe Diarize is a speech-to-text model, designed for accurate audio transcription and recognition. | N/A | N/A | 0 |
ImageImage EditChatAsyncAn image generation model by Black Forest Labs optimized for context-aware and instruction-based image editing. | N/A | N/A | 0 |
FilesVisionOpenAI ChatGPT Image is an image generation capability within ChatGPT, allowing users to create and refine images directly from conversational text prompts. | N/A | N/A | 0 |
ReasoningToolsFilesVisionOpenAI GPT-5 Codex Mini is a code-specialized variant in the GPT-5 series, optimized for code generation, debugging, and software development tasks. | N/A | N/A | 0 |
ChatImageImage EditOpenAI GPT-4o Image is an image generation model in the GPT-4o family, creating visuals from text prompts with fast inference for chat, design, and multimodal agent use cases. | N/A | N/A | 0 |
ChatMiniMax M2.5 Lightning is the lowest-latency MiniMax M2.5 tier for voice-first assistants, live streaming chat, and interactive entertainment. | N/A | N/A | 0 |
ChatWeb SearchOpenAI GPT-4 Net is a language model in the GPT-4 series, offering general-purpose reasoning, code generation, and multimodal capabilities. | N/A | N/A | 0 |
FilesVisionImageChatOpenAI GPT Image 1.5 is an image generation model that produces detailed visuals from text prompts, supporting creative workflows and multimodal applications on OpenAI APIs. | N/A | N/A | 0 |
Alibaba Qwen3 TTS Flash Realtime is a text-to-speech model, designed for high-quality speech synthesis from text input. | N/A | N/A | 0 |
ChatGPTsOpenAI GPT-4 Gizmo is an instruction-tuned variant in the GPT-4 series, optimized for following instructions and conversational tasks. | N/A | N/A | 0 |
MiniMax Feed is a MiniMax platform API endpoint for feed-based content delivery and streaming interactions, exposed by third-party OpenAI-compatible gateways alongside other MiniMax multimodal services. | N/A | N/A | 0 |
Net GPT-4oFree OpenAI Net GPT-4o is a language model in the GPT series, offering general-purpose reasoning, code generation, and multimodal capabilities. | N/A | N/A | 0 |
MiniMax-Text-01Free MiniMax-Text-01 is MiniMax's open-weight flagship language model with hybrid Lightning Attention and MoE architecture, supporting up to 4M tokens of context for long-document reasoning, coding, and agent workflows. | N/A | N/A | 0 |
Zhipu GLM-4 Long extends GLM-4 with extra-long context windows for document QA, retrieval, and knowledge workflows. | N/A | N/A | 0 |
文本Anthropic Claude Instant 1.2 is a legacy fast Claude model for low-latency chat, classification, and lightweight text tasks in older API integrations. | N/A | N/A | 0 |
ToolsVision128KChatOpenAI ChatGPT-4o is the GPT-4o model exposed through ChatGPT endpoints, offering multimodal chat, vision, and fast conversational responses. | 100.72 t/s | 3.47 s | 2 |
ReasoningTools131.1KReasoningGrok 3 Mini is a compact language model in the Grok series, optimized for low-latency responses and efficient inference. | N/A | N/A | 0 |
Alibaba Qwen3.5 Plus Image is a text-to-image generation model in the Qwen 3.5 series, producing high-resolution images from text prompts with strong instruction following and bilingual text rendering. | N/A | N/A | 0 |
文本Anthropic Claude 2.1 is a language model in the Claude series, offering general-purpose reasoning, code generation, and multimodal capabilities. | N/A | N/A | 0 |
OpenAI GPT-4o Mini TTS is a text-to-speech model, designed for high-quality speech synthesis from text input. | N/A | N/A | 0 |
Zhipu AI GLM-4.5 AirX is an enhanced GLM-4.5 Air build with higher throughput for enterprise copilots and bilingual knowledge retrieval. | N/A | N/A | 0 |
ReasoningToolsFilesVisionGoogle Gemini Flash Lite is a compact language model in the Gemini series, optimized for low-latency responses and efficient inference. | N/A | N/A | 0 |
ChatVisionDeprecatedOpenAI GPT-4 Vision is a multimodal vision-language model in the GPT-4 series, supporting both text and image understanding. | N/A | N/A | 0 |
Google Gemini 1.5 Pro 001 is a high-capability language model in the Gemini series, offering enhanced reasoning, code generation, and multimodal capabilities. | N/A | N/A | 0 |
GPT-SoVITS is an open text-to-speech model that synthesizes high-quality speech from text with few-shot voice adaptation, popular for dubbing, localization, and voice cloning projects. | N/A | N/A | 0 |
Alibaba Qwen3 S2S Flash Realtime is a realtime audio model in the Qwen series, supporting low-latency speech and conversational interactions. | N/A | N/A | 0 |
OpenAI GPT-4.5 advances general reasoning, tool use, and multimodal understanding for demanding assistant and agent applications. | N/A | N/A | 0 |
ReasoningToolsFilesVisionAnthropic Claude 3.5 Sonnet set a strong price-performance benchmark for coding, analysis, and vision-enabled assistants before the Claude 4 generation. | N/A | N/A | 0 |
Alibaba Qwen3 Omni Flash is a multimodal model in the Qwen series, offering text, image, video, and audio understanding capabilities. | N/A | N/A | 0 |
32.8KToolsOpen Weights128KMistral AI Devstral Small is a code-specialized variant in the Mistral series, optimized for code generation, debugging, and software development tasks. | N/A | N/A | 0 |
ToolsFilesVisionAudioGoogle Gemini 1.5 Pro is a high-capability language model in the Gemini series, offering enhanced reasoning, code generation, and multimodal capabilities. | N/A | N/A | 0 |
ReasoningToolsFilesVisionGoogle Gemini 3 Pro is a high-capability language model in the Gemini series, offering enhanced reasoning, code generation, and multimodal capabilities. | 12.54 t/s | 3.36 s | 5 |
Reasoning128KChatReasoningOpenAI O1 Mini is a reasoning model in the O series, designed for complex reasoning, problem-solving, and analytical tasks. | N/A | N/A | 0 |
VideoAsyncMiniMax Hailuo 2.3 is a video and multimodal generation model in the Hailuo series, designed for text-to-video and creative media workflows. | N/A | N/A | 0 |
ReasoningTools200KOpen WeightsZhipu GLM-5.1 is a next-generation GLM model aimed at frontier reasoning, coding, and bilingual agent applications. | Dext APIZero APIMoyanjdc API +46 moreShow fewerFuture Hub霸气公益平台KternaDMXAPIAI Claw APIEnenCloud API933999 APIFeng Love APILLM API枫叶17NAS API情酱的API站Mitchll-APIDawnLoadAI DF2ThatAPI老魔公益站百万APIAIO通用智能服务平台PICO AI42公益站DeepRouter随时跑路公益站Neb 公益站Koyeb AI GatewaySWT-APIZetaTechs APISmart APIOptAITommyLam API毫秒APIDibin84 API HubSeamee APISoul 公益站wuer的api站Mars HKMineWuer API星见雅 API6i2MapleLeaf APIAI APIWAADRIMyNav AIC85 API温云OpenOpen8 APIAWA1 API | 41.70 t/s | 11.00 s | 20 |
ReasoningToolsFilesVisionAlibaba Qwen3.6 Plus is an enhanced Qwen3.6-tier model optimized for reasoning, coding, and long-context tasks with balanced cost and performance. | 51.71 t/s | 16.15 s | 5 |
ReasoningToolsFilesVisionZhipu AI GLM-5V Turbo is a multimodal vision-language model in the GLM series, supporting both text and image understanding. | N/A | N/A | 0 |
A model–provider pair whose current input and output prices are both $0 per token. Some providers offer permanent free tiers, others only give one-time credits to new accounts. We mark an offering as free only while its public pricing is zero, and re-check pricing pages regularly.
Most have rate limits — requests per minute, daily caps, or context-length limits — and many require account signup with a verified phone or payment method. Some are time-limited promotions. Always read the provider's terms and quotas before depending on a free endpoint.
Some entries are community-run relays ("公益站") that bundle paid upstream keys and redistribute access for free at the operator's expense. They often advertise larger quotas and a broader model list than official free tiers, but reliability is much lower: operators can pull the plug or disappear overnight, pricing and quotas can change without notice, and many sites are invite-only — requiring a GitHub invite, a forum referral, or a closed community to register. Some keep signups disabled indefinitely. Treat them as best-effort backup channels; keep anything important on official paid endpoints.
Speed varies by model and provider. Sort the table by Most tested for the most reliable benchmarks, or pick a model family from the chips above to drill in. Each row shows median tokens per second and first-token latency from real API tests.
We send identical prompts to each provider through a five-round stress test, count output tokens with tiktoken, and measure both throughput (tokens per second) and time to first token. Numbers are aggregated as medians to resist outliers and refresh on a regular cadence.
For prototypes, side projects, and low-traffic tools, yes. Production traffic will usually hit a rate limit quickly. Treat the free tier as an evaluation channel: validate the model and provider, then move to a paid endpoint with the same model when you scale.
Either no provider currently offers it for free, the free promotion ended, or it has not been benchmarked yet. Open the model's main page to compare paid options, or let us know about a missing free provider via the feedback link in the footer.