Qwen Plus 0728 (thinking) API Benchmarks, Pricing & Provider Data
Compare Qwen Plus 0728 (thinking) with another model
Choose a model to open its comparison page.
Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and cost combination.
Specifications
Input and output token limits for this model, plus how it ranks on long-context understanding.
Features
Technical Details
- Input
- Output
- Released
- Sep 2025
- Knowledge cutoff
- 2025-03-31
- Tokenizer
- Qwen3
- Architecture
- text->text
- Moderated
- No
- Supported parameters
- frequency_penaltyinclude_reasoninglogprobsmax_tokenspresence_penaltyreasoningresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Alternatives & Similar Models
GPT-5.4
gpt-5-4
OpenAI GPT-5.4 extends the GPT-5 family with stronger instruction following, deeper tool use, and improved performance on coding, math, and long-document analysis.
GPT-5.3 Codex
gpt-5-3-codex
OpenAI GPT-5.3 Codex is a code-specialized variant in the GPT-5 series, optimized for code generation, debugging, and software development tasks.
GPT-5.2
gpt-5-2
OpenAI GPT-5.2 is a GPT-5 series model emphasizing advanced reasoning, multimodal understanding, and high-quality outputs for complex enterprise workloads.
GPT-5.4 Mini
gpt-5-4-mini
OpenAI GPT-5.4 Mini is a compact language model in the GPT-5 series, optimized for quick responses and high throughput.
Claude Sonnet 4.6
claude-sonnet-4-6
Anthropic Claude Sonnet 4.6 extends the Sonnet line with improved tool use, coding reliability, and long-context performance for everyday production workloads.
Gemini 3 Flash
gemini-3-flash
Google Gemini 3 Flash is a next-generation fast multimodal model for responsive assistants, document understanding, and high-throughput API traffic.
