Qwen
·Released on Feb 25, 2026

Qwen3.5-Flash API Benchmarks, Pricing & Provider Data

Compare Qwen3.5-Flash with another model

Choose a model to open its comparison page.

Share on X
LLM

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference effic...

Specifications

Input and output token limits for this model, plus how it ranks on long-context understanding.

INPUT
1Mtokens
1.2K pages of text
OUTPUT
65.5Ktokens
8K128K1M4M
1M

Features

Technical Details

Input
Output
Released
Feb 2026
Documentation
Tokenizer
Qwen3
Architecture
text+image+video->text
Moderated
No
Supported parameters
frequency_penaltyinclude_reasoningmax_tokenspresence_penaltyreasoningresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_p

OpenRouter endpoints

1 endpoints

Third-party OpenRouter endpoint data, shown separately from LMSpeed measurements. Some 30-minute live performance fields only appear after syncing with an OpenRouter API key.

Provider endpointInputOutput1d uptime30m latency30m throughputContext / output
Alibaba
alibaba
$0.065/M$0.260/M100.0%undefined tokens / undefined tokens

Rankings are based on community-submitted tests and periodic health probes. Advisory only, not official data.Standard benchmark data may include BenchLM and other public sources.

Build with LMSpeed Data

Free

Free public API for LLM pricing, benchmarks & provider data

  • Real-time pricing & availability
  • Speed & latency benchmarks
  • 300+ models, 600+ providers
  • 1,000 requests/day
View API Documentation