DeepRouter appears to operate an OpenAI-compatible API gateway at deeprouter.top. The public root serves a New API panel describing a unified AI model aggregation and distribution gateway, and its /v1/models endpoint returns an Invalid token new_api_error response.
DeepRouter
deeprouter.top
DeepRouter appears to operate an OpenAI-compatible API gateway at deeprouter.top using the New API panel.
- Models
- 231 models
- From
- --
- Speed
- 92 tok/s
- Updated
- 6/19/2026
- Latency
- 0.00 s
- Created At
- 1/8/2026
- Recharge Rate
- ¥1.00 per $1 quota
Features
Login Methods
API Endpoints
- Blank API Address
https://k.hdgsb.comNo frontend interface
- Overseas CDN Route
https://www.henapi.topOverseas CDN load optimization
- Hong Kong Route
https://deeprouter.topHigh-speed direct connection within China
- Global Direct Routes
https://www.deeprouter.topHigh-quality direct routes worldwide
- US Route
https://us.deeprouter.topDirect connection to US West
Verify ownership to unlock provider management features:
- Edit provider name, content, and links
- Get featured with priority traffic and visibility boost
- Display a verified badge to build user trust
Leaderboard Rankings
Rankings are based on community-submitted tests and periodic health probes. Advisory only, not official data.
About DeepRouter
Health Check
API Benchmarks & Pricing
Compare 172 model rows across audit recency, latest speed tests, throughput, latency, and per-token pricing.
gpt-image-2default | $0.150/M | $0.150/M |
gpt-image-1default | $5.00/M | $40.00/M |
| $0.300/M | $0.300/M | |
gemini-3-pro-imagedefault | $0.300/M | $0.300/M |
gemini-3-pro-image-previewdefault | $0.300/M | $0.300/M |
| $0.400/M | $0.400/M | |
| $0.600/M | $0.600/M | |
cs-claude-opus-4-5default | $0.200/M | $0.200/M |
cs-claude-opus-4-5-thinkingdefault | $0.200/M | $0.200/M |
cs-claude-opus-4-6default | $0.200/M | $0.200/M |
cs-claude-opus-4-6-thinkingdefault | $0.200/M | $0.200/M |
cs-claude-opus-4-7default | $0.200/M | $0.200/M |
cs-claude-opus-4-7-thinkingdefault | $0.200/M | $0.200/M |
cs-claude-sonnet-4-6default | $0.150/M | $0.150/M |
cs-claude-sonnet-4-6-thinkingdefault | $0.075/M | $0.075/M |
cs-gemini-2.5-prodefault | $0.050/M | $0.050/M |
cs-gemini-3-prodefault | $0.100/M | $0.100/M |
cs-gemini-3.1-prodefault | $0.100/M | $0.100/M |
grok-4.2default | $2.00/M | $6.00/M |
grok-4.2-imagedefault | $0.060/M | $0.060/M |
mimo-v2-prodefault | $7.00/M | $21.00/M |
claude-opus-4-7default | $5.00/M | $25.00/M |
| $0.250/M | $1.50/M | |
whisper-1default | $30.00/M | $30.00/M |
gpt-image-2-prodefault | $0.400/M | $0.400/M |
Showing 25 of 172 model rows
Recent Test Records
| Time | Model | Speed | Latency |
|---|---|---|---|
| Jan 8, 06:58 AM | gemini-2.5-pro | 91.91 tok/s | 16.35s |
Similar API Provider Alternatives to Compare
Compare DeepRouter alternatives against 6 nearby API providers using 771 LMSpeed signals across shared model coverage, pricing, benchmark speed, uptime, and free-model availability.
| Provider | Why compare | Models | Free | Avg price | Speed | 30d uptime |
|---|---|---|---|---|---|---|
| DeepRouter deeprouter DeepRouter appears to operate an OpenAI-compatible API gateway at deeprouter.top using the New API panel. | Current provider baseline | 231 | 173 | N/A | 92 tok/s | 9970% |
| Koyeb AI Gateway new-api-koyeb-app An OpenAI-compatible API gateway deployed on Koyeb, providing access to multiple AI models. |
| 578 | 420 | N/A | 33 tok/s | 9950% |
| APIMart apimart APIMart is a pay-as-you-go, OpenAI-compatible AI API gateway operated by Hangzhou Huanzhi Network Technology Co., Ltd. It provides unified access to more than 500 third-party text, image, video, and audio models, with consolidated billing, multi-provider routing, and automatic failover. |
| 257 | 16 | $0.015/M | N/A | 9980% |
| Fengsili API api-fengsili-online An OpenAI-compatible API relay service providing access to multiple AI models. |
| 380 | 0 | $3.00/M | 73 tok/s | 4860% |
| DeadlySignal API deadlysignal |
| 290 | 0 | N/A | 63 tok/s | 9960% |
| 91VIP API hcg-pippi-top A unified API gateway providing access to multiple large language models and AI services with competitive pricing and stability. |
| 181 | 30 | $0.467/M | 148 tok/s | 9890% |
| 柠檬API new-lemonapi-site An AI model API aggregation platform providing unified access to multiple large language models for streamlined integration. |
| 148 | 0 | $0.770/M | 109 tok/s | 2120% |
Announcements
grok-4.6 is now supported. It is recommended to use the Codex group and client access.
WeChat login is temporarily restored. Please bind your email on your profile page as soon as possible, then reset your password on the login page and log in with email. WeChat login will be completely discontinued soon!
Vercel group multiplier adjusted to 0.3. Claude Code ultra-low-cost reverse channel, cache is low and unstable.
Latest support for gemini-3.5-flash
Now supports OpenAI drawing model gpt-image-2
claude-opus-4-7 model is now available; upgrade CC to the latest version to use it.
FAQ
Why does my bill show multiple records when I only sent one message?
This is normal. Many Agent clients do not follow a "one user message = one model request" pattern. For example, Agent tools like Claude CLI or Codex Desktop may make multiple model requests in succession to complete a single task: - Understand the request - Read context - Call tools - Summarize results - Plan next steps Additionally, some clients automatically initiate auxiliary requests even without you sending a new message, such as: - Generating homepage recommendations - Loading or summarizing context - Checking session state - Preprocessing for next steps The billing details show the **actual underlying model requests**, so you may see multiple records around the same time. ### Simple Explanation > **You see one operation or conversation; the system records all actual model calls happening in the background.**
Why does the bill show gpt-5.4-mini when I selected gpt-5.5?
Agent tools may break a task into multiple steps, using different models for different steps. For example, one step might use a stronger model for planning, while another uses a lighter model for simple content processing. So seeing multiple models in the same task does not necessarily indicate an anomaly; it's the Agent client calling different models for different steps.
What are Cache reads and Cache writes? Why is the amount difference large?
Cache is the context caching mechanism of the model service. In consecutive multi-turn tasks, Agents often repeatedly carry large amounts of context, such as project files, conversation history, tool results, etc. - Cache write: Writing context to cache for the first time, usually more expensive. - Cache read: Reusing cached context later, usually much cheaper than re-input. So you may see some records with high Cache writes and higher costs, and others with high Cache reads and lower costs. This usually indicates the Agent is reusing context cache.
Why do some requests have few input tokens but still cost a lot?
Because model billing is not only based on regular input and output tokens, but may also include Cache writes, Cache reads, model price differences, and other factors. Especially in Claude / Agent scenarios, a large portion of the cost may come from context caching, not just the few sentences you manually typed.
Why is the amount to be deducted different from the actual amount deducted?
"Amount to be deducted" can be understood as a reference amount calculated based on the model's official or standard price; "Actual amount deducted" is the amount actually deducted from your balance after applying the current usage group discount. If the page shows labels like "2.8折" (28% off), "6折" (60% off), "1.1折" (11% off), etc., it means the request received the corresponding discount.
Why does it feel like consumption is accelerating? How can I reduce quota usage?
- In the same session, as you continue conversing, the client typically sends more historical messages, file content, tool results, etc., to the model, making the context of each request longer. This makes each request feel more expensive. It's recommended to start a new session after a long conversation, compress context, or choose a lighter model for simple questions to reduce per-request cost. - If you frequently switch between completely different tasks in one session (e.g., analyzing code, then processing images, then writing copy), new questions introduce different contexts, causing previously reusable caches to be overwritten or have lower hit rates. It's recommended to separate different task types into different sessions (e.g., code analysis, image processing, copywriting each in their own session) to improve cache reuse.
When should I contact customer support?
Please contact us for investigation if you encounter the following: - Billing records are continuously generated even when you are clearly not initiating any tasks. - Repeated Cache writes in a short period with no Cache reads. - Abnormally high duplicate charges for the same request. - The model in the billing record completely does not match your usage scenario. - The actual amount deducted is clearly abnormal. - Agent calls fail but still generate incomprehensible charge records. Please join the group and contact customer support, providing the request time, model name, source, and screenshots if possible, to help us quickly locate the issue.
How to troubleshoot 401, 403, 404, and network errors?
- 401: Missing or invalid API key, or authentication source conflict. - 403: Insufficient balance, token group, model permissions, or policy rejection. - 404: Base URL, /v1/responses path, or model ID mismatch. - 429: Reduce concurrency, back off and retry as per server hints. - 5xx: Keep the request ID, briefly back off and retry once. - Connection failed: First verify the same URL with curl; then check DNS, proxy, certificates, and IDE startup environment. Do not directly conclude it's an MTU issue just because other terminal tools can connect. Only adjust network parameters after packet capture or reproducible packet size tests prove it.
Notes
- Health checks: Scope: the 72-hour chart and recent availability measure API connectivity only. Each bar summarizes one hour of checks. Targets: LMSpeed tries the configured health check URL and provider status URL first, then API endpoints derived from known API hosts and recent speed-test base URLs. A website host is considered only when it looks like an API endpoint. Probe steps: each candidate goes through DNS lookup, TCP connection, TLS handshake for HTTPS, and an HTTP HEAD request with redirects followed. Probing stops after the first reachable candidate. Reachable criteria: every required network step must succeed. An HTTP response below 500 is treated as reachable, including 401 because it confirms that an authenticated API endpoint responded, except for statuses classified as blocked. Blocked results: HTTP 403, 429, 521, 525, and 530, plus detected WAF or Cloudflare challenges, are shown as blocked and excluded from availability calculations because LMSpeed cannot determine whether the API itself is down. Model availability: when a dedicated test key is configured, LMSpeed sends an authenticated GET request to a derived /models endpoint and compares returned model IDs with this provider's listed models. These per-model results appear in Models & Pricing and are not included in the provider connectivity percentage. Timeouts: TCP connection, TLS handshake, HTTP connectivity, and model requests each use a 20-second timeout. A full run can take longer when several candidates are tried. Frequency: a background worker checks all providers every 5 minutes by default. The 72-hour chart combines those samples into hourly bars, and the schedule may be changed by the service operator. Limit: automated samples are not an SLA and do not guarantee account quota, every model, every region, or successful completion requests. Check the provider's own status page before making operational decisions.
- Domain Rating data is sourced from Ahrefs. It is a 0–100 backlink-based domain strength signal and does not measure API speed or reliability.
- Announcements and FAQ are read from this provider's NewAPI status snapshot when available. LMSpeed stores the original content and optional English translations from the provider status source, then shows the localized fields on this page.

