Data points: 58
Model compare
The readout for GPT-OSS and Kimi K2 Instruct, before the detailed comparison sheet.
Weighted outcome: GPT-OSS. Benchmark capability categories carry 80%, while price, API performance, and availability carry 20%.
Decision read
GPT-OSS
GPT-OSS has the higher weighted result; Model A / B score 80 to 20.
Evidence depth
58 data points
Includes 0 benchmark rows, 0 audit samples, and 7 provider examples.
Selection signal
Start with GPT-OSS
The charts below split 7 high-signal samples across speed, scores, and audit health.
Switch either side of this report to compare another model with the same LMSpeed data pipeline.
Select a different model to open a new comparison URL.
This report only uses LMSpeed data for GPT-OSS and Kimi K2 Instruct: pricing, speed aggregates, third-party benchmark scores, and shared provider samples.
| Model compare | GPT-OSS | Kimi K2 Instruct |
|---|---|---|
| Overall leader | Leading | Contender |
| Weighted overall score | 80.0 pts | 20.0 pts |
| Benchmark category leads | 0 categories | 0 categories |
| Operational advantages | Cheapest input price, Average speed, Free providers, Provider coverage | First-token latency |
| Context window | 131.1K tokens | No data |
| Max output | 131.1K tokens | No data |
| Modalities | Input Text Output Text | No data |
| Features |
The overall result weights benchmark capability categories at 80% and price, API speed/latency, and availability at 20%. Recent test volume does not affect the winner, and missing benchmark categories are excluded.
| Model compare | GPT-OSS | Kimi K2 Instruct |
|---|---|---|
| Developer | No data | Moonshot AI |
| Released | Aug 2025 | No data |
| Parameters | 120B | No data |
| Tokenizer | GPT | No data |
| Knowledge cutoff | 2024-06-30 | No data |
| OpenRouter ID | openai/gpt-oss-120b:free | No data |
| References | No data | No data |
This report only uses LMSpeed data for GPT-OSS and Kimi K2 Instruct: pricing, speed aggregates, third-party benchmark scores, and shared provider samples.
GPT-OSS
GPT-OSS has these operational advantages: Cheapest input price, Average speed, Free providers, Provider coverage.
Kimi K2 Instruct
Kimi K2 Instruct has these operational advantages: First-token latency.
Third-party benchmark profile synced into LMSpeed; only metrics available for both models are shown.
Metric-level scores with benchmark source, rank depth, confidence, error, and evaluation date where available.
No shared professional benchmark scores are available yet.
Latest completed audits from shared providers, with four safety and integrity score groups plus report links.
| Provider | GPT-OSS | Kimi K2 Instruct |
|---|---|---|
| No completed audits are available from shared providers yet. | ||
Speed aggregates and input/output pricing share each provider row for real API selection and migration cost checks.
| Provider | GPT-OSS | Kimi K2 Instruct |
|---|---|---|
120 tests | GPT-OSS speed / latency 153 tok/s / 1011ms input / output No data | Kimi K2 Instruct speed / latency 45 tok/s / 578ms input / output No data |
70 tests | GPT-OSS speed / latency 198 tok/s / 7216ms input / output No data | Kimi K2 Instruct speed / latency N/A / N/A input / output No data |
70 tests | GPT-OSS speed / latency N/A / N/A input / output No data | Kimi K2 Instruct speed / latency 182 tok/s / 1382ms input / output No data |
35 tests | GPT-OSS speed / latency 1061 tok/s / 857ms input / output No data | Kimi K2 Instruct speed / latency N/A / N/A input / output No data |
35 tests | GPT-OSS speed / latency 347 tok/s / 2670ms input / output No data | Kimi K2 Instruct speed / latency N/A / N/A input / output No data |
GPT-OSS gpt-oss-20b speed / latency No data input / output $0/M/$0/M | Kimi K2 Instruct kimi-k2-instruct speed / latency No data input / output $0.038/M | |
GPT-OSS openai/gpt-oss-20b speed / latency No data input / output $0/request | Kimi K2 Instruct kimi-k2-instruct speed / latency No data input / output $0.100/request |
Weighted outcome: GPT-OSS. Benchmark capability categories carry 80%, while price, API performance, and availability carry 20%.
Text inputText outputTool callingReasoning |
| None listed |