Choose a model to open its comparison page.
Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion tot...
Observed-capability estimates with uncertainty reported separately from independent benchmark families.
Input and output token limits for this model, plus how it ranks on long-context understanding.
2 metrics
2 metrics
0 metrics · No data
1 metric · Provisional
2 metrics · Provisional
0 metrics · No data
0 metrics · No data
0 metrics · No data
0 metrics · No data
0 metrics · No data
Third-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.
| Provider endpoint | Input | Output | 1d uptime | 30m latency | 30m throughput | Context / output |
|---|---|---|---|---|---|---|
Alibaba alibaba | $2/M | $6/M | 100.0% | — | — | 1M tokens / 131.1K tokens |
DigitalOcean digitalocean | $2/M | $6/M | 99.7% | — | — | 262.1K tokens / 52.4K tokens |
SiliconFlow siliconflow/fp8 | $2/M | $6/M | 99.2% | — | — | 1.0M tokens / 131.1K tokens |
Together together | $2.50/M | $6.25/M | 98.2% | — | — | 1.0M tokens / — |
Modal modal/nvfp4 | $2/M | $6/M | 98.2% | — | — | 1M tokens / 262.1K tokens |
Venice venice | $2.50/M | $7.50/M | 95.5% | — | — | 262.1K tokens / 65.5K tokens |
DeepInfra deepinfra/fp4 | $2/M | $6/M | 94.5% | — | — | 262.1K tokens / 131.1K tokens |