Gemma 3 12B API Benchmarks, Pricing & Provider Data
Compare Gemma 3 12B with another model
Choose a model to open its comparison page.
Gemma 3 12B API pricing covers undefined API provider} other undefined API providers}}, from $0.000001/request to $0.0055/M.
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, ...
Specifications
Input and output token limits for this model, plus how it ranks on long-context understanding.
Features
Technical Details
- Input
- Output
- Total parameters
- 12B
- Released
- Mar 2025
- Knowledge cutoff
- 2024-08-31
- Tokenizer
- Gemini
- Architecture
- text+image->text
- Instruct type
- gemma
- Moderated
- No
- Supported parameters
- frequency_penaltylogit_biasmax_tokensmin_ppresence_penaltyrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_p
OpenRouter endpoints
1 endpointsThird-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.
| Provider endpoint | Input | Output | 1d uptime | 30m latency | 30m throughput | Context / output |
|---|---|---|---|---|---|---|
DeepInfra deepinfra/bf16 | $0.050/M | $0.150/M | 100.0% | — | — | undefined tokens / undefined tokens |
Pricing Comparison
Compare Gemma 3 12B API pricing across 2 providers. Prices range from $0.000001/request to $0.0055/M. LLMService offers the lowest rate at $0.000001/request.
| Provider | Health | Model Variant | Group | Input ($/M) | Output ($/M) | Speed (t/s) | First token | Audit |
|---|---|---|---|---|---|---|---|---|
L1 99% | google/gemma-3-12b-it:free | default | $0.000001/request | - | — | — | — | |
L1 0% | gemma-3-12b-it | Gemini | $0.0055/M | $0.011/M | — | — | — |
Alternatives & Similar Models
GPT-5.4
gpt-5-4
OpenAI GPT-5.4 extends the GPT-5 family with stronger instruction following, deeper tool use, and improved performance on coding, math, and long-document analysis.
GPT-5.3 Codex
gpt-5-3-codex
OpenAI GPT-5.3 Codex is a code-specialized variant in the GPT-5 series, optimized for code generation, debugging, and software development tasks.
GPT-5.4 Mini
gpt-5-4-mini
OpenAI GPT-5.4 Mini is a compact language model in the GPT-5 series, optimized for quick responses and high throughput.
Kimi K2.5
kimi-k2-5
Moonshot Kimi K2.5 is an open-weight multimodal agent model with native vision and text input, strong coding performance, and a 256K context window.
MiniMax M2.5
minimax-m2-5
MiniMax M2.5 is MiniMax's flagship text model for coding and agents, with SOTA-level programming and agentic performance, improved token efficiency, and fast high-TPS API deployment.
DeepSeek V3.2
deepseek-v3-2
DeepSeek V3.2 is an upgraded V3-series MoE model with stronger reasoning, coding, and math performance, widely available through OpenAI-compatible API relays.
Frequently Asked Questions
- What benchmark data does Gemma 3 12B include?
- LMSpeed shows Gemma 3 12B benchmark context, API price, output speed, first-token latency, and provider data across 2 providers when those signals are available.
- What is the Gemma 3 12B API price?
- Gemma 3 12B has pricing from undefined provider} other undefined providers}}, ranging from $0.000001/request to $0.0055/M. LLMService has the lowest listed price.
- What does the Gemma 3 12B API pricing table include?
- The Gemma 3 12B API pricing table compares 2 providers by input price, output price, free tier status, speed, first-token latency, and recent health data when available.
- Which provider has the cheapest Gemma 3 12B API pricing?
- LLMService currently has the lowest listed Gemma 3 12B price at $0.000001/request across undefined provider} other undefined providers}}.
- Is Gemma 3 12B API free?
- Gemma 3 12B does not currently have a free API tier on LMSpeed. All 2 providers charge per token.
