Gemma 3 4B API Benchmarks, Pricing & Provider Data
Compare Gemma 3 4B with another model
Choose a model to open its comparison page.
Gemma 3 4B API pricing covers undefined API provider} other undefined API providers}}, from $0.000001/request to $0.011/M.
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, ...
Specifications
Input and output token limits for this model, plus how it ranks on long-context understanding.
Features
Technical Details
- Input
- Output
- Total parameters
- 4B
- Released
- Mar 2025
- Knowledge cutoff
- 2024-08-31
- Tokenizer
- Gemini
- Architecture
- text+image->text
- Instruct type
- gemma
- Moderated
- No
- Supported parameters
- frequency_penaltylogit_biasmax_tokensmin_ppresence_penaltyrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetop_ktop_p
OpenRouter endpoints
1 endpointsThird-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.
| Provider endpoint | Input | Output | 1d uptime | 30m latency | 30m throughput | Context / output |
|---|---|---|---|---|---|---|
DeepInfra deepinfra/bf16 | $0.050/M | $0.100/M | 100% | — | — | undefined tokens / undefined tokens |
Pricing Comparison
Compare Gemma 3 4B API pricing across 2 providers. Prices range from $0.000001/request to $0.011/M. LLMService offers the lowest rate at $0.000001/request.
| Provider | Health | Model Variant | Group | Input ($/M) | Output ($/M) | Speed (t/s) | First token | Audit |
|---|---|---|---|---|---|---|---|---|
L1 99% | google/gemma-3-4b-it:free | default | $0.000001/request | - | — | — | — | |
L1 0% | gemma-3-4b-it | Gemini | $0.011/M | $0.011/M | — | — | — |
Alternatives & Similar Models
GPT-5.4
gpt-5-4
OpenAI GPT-5.4 extends the GPT-5 family with stronger instruction following, deeper tool use, and improved performance on coding, math, and long-document analysis.
GPT-5.3 Codex
gpt-5-3-codex
OpenAI GPT-5.3 Codex is a code-specialized variant in the GPT-5 series, optimized for code generation, debugging, and software development tasks.
GPT-5.4 Mini
gpt-5-4-mini
OpenAI GPT-5.4 Mini is a compact language model in the GPT-5 series, optimized for quick responses and high throughput.
Kimi K2.5
kimi-k2-5
Moonshot Kimi K2.5 is an open-weight multimodal agent model with native vision and text input, strong coding performance, and a 256K context window.
MiniMax M2.5
minimax-m2-5
MiniMax M2.5 is MiniMax's flagship text model for coding and agents, with SOTA-level programming and agentic performance, improved token efficiency, and fast high-TPS API deployment.
DeepSeek V3.2
deepseek-v3-2
DeepSeek V3.2 is an upgraded V3-series MoE model with stronger reasoning, coding, and math performance, widely available through OpenAI-compatible API relays.
Frequently Asked Questions
- What benchmark data does Gemma 3 4B include?
- LMSpeed shows Gemma 3 4B benchmark context, API price, output speed, first-token latency, and provider data across 2 providers when those signals are available.
- What is the Gemma 3 4B API price?
- Gemma 3 4B has pricing from undefined provider} other undefined providers}}, ranging from $0.000001/request to $0.011/M. LLMService has the lowest listed price.
- What does the Gemma 3 4B API pricing table include?
- The Gemma 3 4B API pricing table compares 2 providers by input price, output price, free tier status, speed, first-token latency, and recent health data when available.
- Which provider has the cheapest Gemma 3 4B API pricing?
- LLMService currently has the lowest listed Gemma 3 4B price at $0.000001/request across undefined provider} other undefined providers}}.
- Is Gemma 3 4B API free?
- Gemma 3 4B does not currently have a free API tier on LMSpeed. All 2 providers charge per token.
