deeprouter.top

DeepRouter appears to operate an OpenAI-compatible API gateway at deeprouter.top using the New API panel.

Models
231 models
From
--
Speed
92 tok/s
Updated
6/19/2026
Latency
0.00 s
Created At
1/8/2026
Recharge Rate
¥1.00 per $1 quota

Features

DrawingTaskData ExportCheck-in

Login Methods

GitHubWeChat

API Endpoints

  • Blank API Address
    https://k.hdgsb.com

    No frontend interface

  • Overseas CDN Route
    https://www.henapi.top

    Overseas CDN load optimization

  • Hong Kong Route
    https://deeprouter.top

    High-speed direct connection within China

  • Global Direct Routes
    https://www.deeprouter.top

    High-quality direct routes worldwide

  • US Route
    https://us.deeprouter.top

    Direct connection to US West

Claim this provider

Verify ownership to unlock provider management features:

  • Edit provider name, content, and links
  • Get featured with priority traffic and visibility boost
  • Display a verified badge to build user trust

Leaderboard Rankings

Rankings are based on community-submitted tests and periodic health probes. Advisory only, not official data.

About DeepRouter

DeepRouter appears to operate an OpenAI-compatible API gateway at deeprouter.top. The public root serves a New API panel describing a unified AI model aggregation and distribution gateway, and its /v1/models endpoint returns an Invalid token new_api_error response.

Health Check

100%Recent availability
History (72 pts)
PastNow

API Benchmarks & Pricing

Compare 172 model rows across audit recency, latest speed tests, throughput, latency, and per-token pricing.

Model
Input ($/M)
Output ($/M)
$0.150/M$0.150/M
$5.00/M$40.00/M
$0.300/M$0.300/M
$0.300/M$0.300/M
$0.300/M$0.300/M
$0.400/M$0.400/M
$0.600/M$0.600/M
cs-claude-opus-4-5default
$0.200/M$0.200/M
cs-claude-opus-4-5-thinkingdefault
$0.200/M$0.200/M
cs-claude-opus-4-6default
$0.200/M$0.200/M
cs-claude-opus-4-6-thinkingdefault
$0.200/M$0.200/M
cs-claude-opus-4-7default
$0.200/M$0.200/M
cs-claude-opus-4-7-thinkingdefault
$0.200/M$0.200/M
cs-claude-sonnet-4-6default
$0.150/M$0.150/M
cs-claude-sonnet-4-6-thinkingdefault
$0.075/M$0.075/M
cs-gemini-2.5-prodefault
$0.050/M$0.050/M
cs-gemini-3-prodefault
$0.100/M$0.100/M
cs-gemini-3.1-prodefault
$0.100/M$0.100/M
grok-4.2default
$2.00/M$6.00/M
grok-4.2-imagedefault
$0.060/M$0.060/M
$7.00/M$21.00/M
$5.00/M$25.00/M
$0.250/M$1.50/M
whisper-1default
$30.00/M$30.00/M
gpt-image-2-prodefault
$0.400/M$0.400/M

Showing 25 of 172 model rows

Recent Test Records

TimeModelSpeedLatency
Jan 8, 06:58 AM
gemini-2.5-pro
91.91 tok/s
16.35s

Similar API Provider Alternatives to Compare

Compare DeepRouter alternatives against 6 nearby API providers using 771 LMSpeed signals across shared model coverage, pricing, benchmark speed, uptime, and free-model availability.

ProviderWhy compareModelsFreeAvg priceSpeed30d uptime
DeepRouter

deeprouter

DeepRouter appears to operate an OpenAI-compatible API gateway at deeprouter.top using the New API panel.

Current provider baseline231173N/A92 tok/s9970%
Koyeb AI Gateway

new-api-koyeb-app

An OpenAI-compatible API gateway deployed on Koyeb, providing access to multiple AI models.

  • More free-model options
  • Broader model coverage
578420N/A33 tok/s9950%
APIMart

apimart

APIMart is a pay-as-you-go, OpenAI-compatible AI API gateway operated by Hangzhou Huanzhi Network Technology Co., Ltd. It provides unified access to more than 500 third-party text, image, video, and audio models, with consolidated billing, multi-provider routing, and automatic failover.

  • Higher 30-day availability
  • Broader model coverage
25716$0.015/MN/A9980%
Fengsili API

api-fengsili-online

An OpenAI-compatible API relay service providing access to multiple AI models.

  • Broader model coverage
3800$3.00/M73 tok/s4860%
DeadlySignal API

deadlysignal

  • Broader model coverage
2900N/A63 tok/s9960%
91VIP API

hcg-pippi-top

A unified API gateway providing access to multiple large language models and AI services with competitive pricing and stability.

  • Faster measured speed
18130$0.467/M148 tok/s9890%
柠檬API

new-lemonapi-site

An AI model API aggregation platform providing unified access to multiple large language models for streamlined integration.

  • Faster measured speed
1480$0.770/M109 tok/s2120%

Announcements

default8/13/2026

grok-4.6 is now supported. It is recommended to use the Codex group and client access.

warning8/12/2026

WeChat login is temporarily restored. Please bind your email on your profile page as soon as possible, then reset your password on the login page and log in with email. WeChat login will be completely discontinued soon!

success6/16/2026

Vercel group multiplier adjusted to 0.3. Claude Code ultra-low-cost reverse channel, cache is low and unstable.

success5/20/2026

Latest support for gemini-3.5-flash

success4/22/2026

Now supports OpenAI drawing model gpt-image-2

success4/17/2026

claude-opus-4-7 model is now available; upgrade CC to the latest version to use it.

FAQ

Why does my bill show multiple records when I only sent one message?

This is normal. Many Agent clients do not follow a "one user message = one model request" pattern. For example, Agent tools like Claude CLI or Codex Desktop may make multiple model requests in succession to complete a single task: - Understand the request - Read context - Call tools - Summarize results - Plan next steps Additionally, some clients automatically initiate auxiliary requests even without you sending a new message, such as: - Generating homepage recommendations - Loading or summarizing context - Checking session state - Preprocessing for next steps The billing details show the **actual underlying model requests**, so you may see multiple records around the same time. ### Simple Explanation > **You see one operation or conversation; the system records all actual model calls happening in the background.**

Why does the bill show gpt-5.4-mini when I selected gpt-5.5?

Agent tools may break a task into multiple steps, using different models for different steps. For example, one step might use a stronger model for planning, while another uses a lighter model for simple content processing. So seeing multiple models in the same task does not necessarily indicate an anomaly; it's the Agent client calling different models for different steps.

What are Cache reads and Cache writes? Why is the amount difference large?

Cache is the context caching mechanism of the model service. In consecutive multi-turn tasks, Agents often repeatedly carry large amounts of context, such as project files, conversation history, tool results, etc. - Cache write: Writing context to cache for the first time, usually more expensive. - Cache read: Reusing cached context later, usually much cheaper than re-input. So you may see some records with high Cache writes and higher costs, and others with high Cache reads and lower costs. This usually indicates the Agent is reusing context cache.

Why do some requests have few input tokens but still cost a lot?

Because model billing is not only based on regular input and output tokens, but may also include Cache writes, Cache reads, model price differences, and other factors. Especially in Claude / Agent scenarios, a large portion of the cost may come from context caching, not just the few sentences you manually typed.

Why is the amount to be deducted different from the actual amount deducted?

"Amount to be deducted" can be understood as a reference amount calculated based on the model's official or standard price; "Actual amount deducted" is the amount actually deducted from your balance after applying the current usage group discount. If the page shows labels like "2.8折" (28% off), "6折" (60% off), "1.1折" (11% off), etc., it means the request received the corresponding discount.

Why does it feel like consumption is accelerating? How can I reduce quota usage?

- In the same session, as you continue conversing, the client typically sends more historical messages, file content, tool results, etc., to the model, making the context of each request longer. This makes each request feel more expensive. It's recommended to start a new session after a long conversation, compress context, or choose a lighter model for simple questions to reduce per-request cost. - If you frequently switch between completely different tasks in one session (e.g., analyzing code, then processing images, then writing copy), new questions introduce different contexts, causing previously reusable caches to be overwritten or have lower hit rates. It's recommended to separate different task types into different sessions (e.g., code analysis, image processing, copywriting each in their own session) to improve cache reuse.

When should I contact customer support?

Please contact us for investigation if you encounter the following: - Billing records are continuously generated even when you are clearly not initiating any tasks. - Repeated Cache writes in a short period with no Cache reads. - Abnormally high duplicate charges for the same request. - The model in the billing record completely does not match your usage scenario. - The actual amount deducted is clearly abnormal. - Agent calls fail but still generate incomprehensible charge records. Please join the group and contact customer support, providing the request time, model name, source, and screenshots if possible, to help us quickly locate the issue.

How to troubleshoot 401, 403, 404, and network errors?

- 401: Missing or invalid API key, or authentication source conflict. - 403: Insufficient balance, token group, model permissions, or policy rejection. - 404: Base URL, /v1/responses path, or model ID mismatch. - 429: Reduce concurrency, back off and retry as per server hints. - 5xx: Keep the request ID, briefly back off and retry once. - Connection failed: First verify the same URL with curl; then check DNS, proxy, certificates, and IDE startup environment. Do not directly conclude it's an MTU issue just because other terminal tools can connect. Only adjust network parameters after packet capture or reproducible packet size tests prove it.

Notes

  • Health checks: Scope: the 72-hour chart and recent availability measure API connectivity only. Each bar summarizes one hour of checks. Targets: LMSpeed tries the configured health check URL and provider status URL first, then API endpoints derived from known API hosts and recent speed-test base URLs. A website host is considered only when it looks like an API endpoint. Probe steps: each candidate goes through DNS lookup, TCP connection, TLS handshake for HTTPS, and an HTTP HEAD request with redirects followed. Probing stops after the first reachable candidate. Reachable criteria: every required network step must succeed. An HTTP response below 500 is treated as reachable, including 401 because it confirms that an authenticated API endpoint responded, except for statuses classified as blocked. Blocked results: HTTP 403, 429, 521, 525, and 530, plus detected WAF or Cloudflare challenges, are shown as blocked and excluded from availability calculations because LMSpeed cannot determine whether the API itself is down. Model availability: when a dedicated test key is configured, LMSpeed sends an authenticated GET request to a derived /models endpoint and compares returned model IDs with this provider's listed models. These per-model results appear in Models & Pricing and are not included in the provider connectivity percentage. Timeouts: TCP connection, TLS handshake, HTTP connectivity, and model requests each use a 20-second timeout. A full run can take longer when several candidates are tried. Frequency: a background worker checks all providers every 5 minutes by default. The 72-hour chart combines those samples into hourly bars, and the schedule may be changed by the service operator. Limit: automated samples are not an SLA and do not guarantee account quota, every model, every region, or successful completion requests. Check the provider's own status page before making operational decisions.
  • Domain Rating data is sourced from Ahrefs. It is a 0–100 backlink-based domain strength signal and does not measure API speed or reliability.
  • Announcements and FAQ are read from this provider's NewAPI status snapshot when available. LMSpeed stores the original content and optional English translations from the provider status source, then shows the localized fields on this page.