Models
32 models
From
$0.058/M
Speed
--
Updated
9/11/2026
Latency
--
Created At
8/13/2026
Recharge Rate
¥6.95 per $1 quota

Features

DrawingTaskData Export

Login Methods

GitHubOIDC

Billing

Payment methods

AlipayCorporate transferWeChat Pay

Billing types

Usage-based
Refunds
Supported
Invoicing
Supported

API Endpoints

  • Endpoint 1Historical / Unverified
    https://yomiapi.com
Claim this provider
WVerified by Wellscomet

Contact

  • 交流群(不定期抽福利)1078436366

Leaderboard Rankings

Rankings are based on community-submitted tests and periodic health probes. Advisory only, not official data.

About Yomi API

Yomi API 是 AI 大模型 API 聚合平台。传统中转站只解决"连得上",Yomi 更进一步——专注解决"用得高效":把多模型选型、比价、性能监控、统一计费整合进一套系统。

稳定接入 GPT、Claude、Gemini、DeepSeek 等 40+ 模型,全系 OpenAI 兼容、零改造接入;模型广场实时比价 + 24h 性能快照,让"哪个最划算/最快/最稳"变成可量化决策。按 token 透明计费、无隐藏扣量,一张账单统管全模型消耗,最低可至官方价一折,即充即用、无月费无最低消费。注册即送7元免费额度,交流群不定期发送福利。

东京高防节点 + Cloudflare 加速,SLA 保障,控制台实时公开延迟/TPS/成功率;提供全天售后,不满意随时可退款,尤其适配企业级用户规模化调用。

Health Check

100%Recent availability
History (72 pts)
PastNow

API Benchmarks & Pricing

Compare 55 model rows across audit recency, latest speed tests, throughput, latency, and per-token pricing.

Model
Input ($/M)
Output ($/M)
Model availability
gpt-image-2-proGPT image专用(仅生图)
$0.020/request-
0%
gpt-image-2.5GPT image专用(仅生图)
$0.013/request-
0%
deepseek-v4.1-flashDeepSeek 原生官方
$0.134/M$0.537/M
0%
gpt-image-2.5-sunburstGPT image专用(仅生图)
$0.025/request-
0%
gpt-image-2.5-flareGPT image专用(仅生图)
$0.020/request-
0%
gpt-6-astraGPT 特惠线路优化
$0.600/M$3.00/M
0%
gpt-6-astraGPT Plus线路
$0.800/M$4.00/M
0%
gpt-6-astraGPT Pro 高性能线路
$1.00/M$5.00/M
0%
gemini-3.8-flashGemini线路
$0.120/M$0.600/M
0%
claude-fable-5-1Claude-MAX(cc专用,仅接cli)
$2.90/M$14.50/M
0%
claude-fable-5-1Claude官方Max线路
$3.10/M$15.50/M
0%
glm-5.3-flash智谱 国模折扣
$0.058/M$0.200/M
0%
deepseek-v4-flash-vision-expDeepSeek 原生官方
$0.134/M$0.537/M
0%
glm-5.3智谱 国模折扣
$0.575/M$2.00/M
0%
gemini-3.7-flashGemini线路
$0.120/M$0.600/M
0%
grok-4.6Grok 线路
$0.320/M$0.960/M
0%
claude-opus-5ClaudePlus线路(不支持fable)
$0.600/M$3.00/M
0%
claude-opus-5Claude-MAX(cc专用,仅接cli)
$1.45/M$7.25/M
0%
claude-opus-5Claude官方Max线路
$1.55/M$7.75/M
0%
gemini-3.6-flashGemini线路
$0.240/M$1.20/M
0%
kimi-k3kimi 官方满血线路
$1.50/M$7.50/M
0%
gpt-5.6-terraGPT 特惠线路优化
$0.120/M$0.720/M
0%
gpt-5.6-terraGPT Plus线路
$0.160/M$0.960/M
0%
gpt-5.6-terraGPT Pro 高性能线路
$0.200/M$1.20/M
0%
gpt-5.6-solGPT 特惠线路优化
$0.300/M$1.80/M
0%

Showing 25 of 55 model rows

Recent Test Records

No test records available

Similar API Provider Alternatives to Compare

Compare Yomi API alternatives against 6 nearby API providers using 255 LMSpeed signals across shared model coverage, pricing, benchmark speed, uptime, and free-model availability.

ProviderWhy compareModelsFreeAvg priceSpeed30d uptime
Yomi API

yomi-api

Current provider baseline320$0.129/MN/A9970%
钠 API

naapi-cc

Na API (naapi.cc) is an OpenAI-compatible LLM API gateway with competitive pricing and stable access to 100+ models from OpenAI, Anthropic, Google, and more.

  • Lower average pricing
  • More free-model options
  • Broader model coverage
  • Same provider category
1,0862$0.082/M649 tok/s9950%
VSLLM

vsllm-com

VSLLM runs a New API-powered AI gateway on vsllm.com for aggregated model access through a single endpoint.

  • Lower average pricing
  • More free-model options
  • Broader model coverage
  • Same provider category
2212$0.056/M174 tok/s9970%
10dian-API

api-10dian-ai-top

A multi-model aggregation platform providing unified, stable, and high-performance AI API services for developers.

  • Lower average pricing
  • More free-model options
  • Broader model coverage
  • Same provider category
1519$0.000003/M130 tok/s0%
初叶🍂Furry API

ai-chuyel-top

A free, community-supported API service providing access to various AI models for developers and users.

  • More free-model options
  • Broader model coverage
  • Same provider category
1,069173$8.50/M171 tok/s3360%
天絮 API

tianxu-api

Tianxu API provides an AI model relay service with multiple access points and stable connectivity.

  • Lower average pricing
  • Broader model coverage
  • Same provider category
8100$0.029/M93 tok/s9920%
速创API

suchuang

Suchuang API provides access to various AI models including text generation, image creation, and video generation through a unified API.

  • Lower average pricing
  • Broader model coverage
  • Same provider category
5210$0.020/M74 tok/s9960%

Announcements

service8/17/2026

DeepSeek has recently adjusted the pricing and billing rules for API models, introducing peak/off-peak billing.

To align with upstream official prices, Yomi API will adjust DeepSeek model prices starting August 17, 2026, covering input, cache hit, and output token billing for DeepSeek V4 series models.

This adjustment only syncs upstream official prices and does not affect other models or services. Specific prices are subject to the Yomi API model plaza and pricing page.

Thank you for your understanding and support.

service8/16/2026

The discount multiplier for Zhipu domestic models has been reduced from 0.6 to 0.35, effective immediately. Existing keys and call methods require no changes. Thank you for your support.

service8/15/2026

Taking GPT-5.6-sol as an example, the correct model name is: gpt-5.6-sol Incorrect examples: gpt5.6, gpt5.6sol, entering API Key in the model name field, etc. Model names must exactly match the Model ID provided by the platform. Do not abbreviate, modify, or guess. The Yomi 'Model Services' page provides accurate model names. Use the names as displayed, copy directly. Model name ≠ API Key, do not confuse them.

service8/6/2026

Due to upstream model service adjustments, the GPT-5.6 Luna model will be temporarily suspended starting August 6, 2026 at 10:00 Beijing time. At that time, the model will be removed from the Yomi API model list, and users will no longer be able to call gpt-5.6-luna. If your application or project is using this model, please complete migration in advance to avoid disruption. Yomi API will continue to monitor model service updates and provide stable, reliable alternatives in a timely manner. Thank you for your understanding and support.

promo7/28/2026

Single recharge of 100 yuan gives 20 yuan bonus, 500 yuan gives 120 yuan, 1000 yuan gives 300 yuan. Bonus balance is credited instantly with no expiration.

update7/27/2026

The help documentation has been completely restructured, adding a quick start guide, media knowledge base, and FAQ quick reference.

promo7/27/2026

Single recharge of 100 yuan or more gets an extra 10% bonus. The event deadline is as shown on the recharge page.

update7/26/2026

This upgrade covers four core nodes: Tokyo, Singapore, Frankfurt, and US East. The gateway has been switched to a multi-instance load balancing architecture.

After the upgrade, P99 latency decreased from 840ms to 490ms, and long-connection disconnection rate dropped by 61%. Existing keys require no adjustment; all traffic has been automatically switched to the new gateway.

update7/20/2026

The announcement center now supports category filtering and multiple display formats. The help documentation now includes full-text search; enter keywords to locate specific content paragraphs.

model7/14/2026

Qwen3 Max and DeepSeek R2 have been integrated into the platform's unified catalog. Qwen3 Max focuses on general conversation and long text, while DeepSeek R2 enhances reasoning and mathematical capabilities. Both support all client integrations, and existing keys require no adjustment.

promo6/30/2026

Mid-year rewards event: Single recharge of 100 gets 10% bonus, 500 gets 20%, 1000 gets 30%. Bonus balance is credited instantly, valid for all models, with no expiration.

update6/22/2026

Three new access points have been added: Frankfurt, Silicon Valley, and São Paulo, bringing the total global nodes to 12. Average access latency for European and South American users has decreased by approximately 35%, with automatic optimal routing and no configuration changes required.

service6/15/2026

From 02:17 to 02:29 today, the US East node experienced network jitter. The gateway's auto-retry mechanism was activated, no failed requests occurred, and service has fully recovered.

update6/8/2026

The console usage page has been fully upgraded: added usage breakdown by model, API key, and time dimension, with CSV export support. Cost attribution is now clear at a glance.

model5/29/2026

Claude Sonnet 4.5 has been fully integrated, with stable performance in coding, tool calling, and long-context scenarios. Platform prices are lower than official catalog prices and support all major clients.

promo5/12/2026

Invite new users to register and recharge, and you get 8% cashback of their recharge amount. Friends also enjoy a 5% discount on their first recharge. Invitation links and cashback details are on the 'Referral Rewards' page in the console.

model4/16/2026

Gemini 2.5 Pro has been integrated into the platform, supporting ultra-long context and multimodal input such as images and audio. It uses the same catalog and billing as existing models.

update4/2/2026

The platform's daily average API calls have exceeded 120 million, with overall availability maintained above 99.95%. Thank you for the trust of every developer. We will continue to improve stability.

update3/18/2026

API keys now support setting monthly quota limits and automatic expiration per key, allowing precise cost control even with shared team accounts.

Notes

  • Health checks: Scope: the 72-hour chart and recent availability measure API connectivity only. Each bar summarizes one hour of checks. Targets: LMSpeed tries the configured health check URL and provider status URL first, then API endpoints derived from known API hosts and recent speed-test base URLs. A website host is considered only when it looks like an API endpoint. Probe steps: each candidate goes through DNS lookup, TCP connection, TLS handshake for HTTPS, and an HTTP HEAD request with redirects followed. Probing stops after the first reachable candidate. Reachable criteria: every required network step must succeed. An HTTP response below 500 is treated as reachable, including 401 because it confirms that an authenticated API endpoint responded, except for statuses classified as blocked. Blocked results: HTTP 403, 429, 521, 525, and 530, plus detected WAF or Cloudflare challenges, are shown as blocked and excluded from availability calculations because LMSpeed cannot determine whether the API itself is down. Model availability: when a dedicated test key is configured, LMSpeed sends an authenticated GET request to a derived /models endpoint and compares returned model IDs with this provider's listed models. These per-model results appear in Models & Pricing and are not included in the provider connectivity percentage. Timeouts: TCP connection, TLS handshake, HTTP connectivity, and model requests each use a 20-second timeout. A full run can take longer when several candidates are tried. Frequency: a background worker checks all providers every 5 minutes by default. The 72-hour chart combines those samples into hourly bars, and the schedule may be changed by the service operator. Limit: automated samples are not an SLA and do not guarantee account quota, every model, every region, or successful completion requests. Check the provider's own status page before making operational decisions.
  • Domain Rating data is sourced from Ahrefs. It is a 0–100 backlink-based domain strength signal and does not measure API speed or reliability.
  • Announcements and FAQ are read from this provider's NewAPI status snapshot when available. LMSpeed stores the original content and optional English translations from the provider status source, then shows the localized fields on this page.