VSLLM runs a New API-powered AI gateway on vsllm.com for aggregated model access through a single endpoint.

Models
209 models
From
$0.110/M
Speed
174 tok/s
Updated
7/2/2026
Latency
0.00 s
Created At
3/8/2026
Recharge Rate
¥7.30 per $1 quota

Features

DrawingTaskData Export

Login Methods

GitHubOIDC

API Endpoints

  • Endpoint 1Historical / Unverified
    https://vsllm.com
  • Endpoint 2Historical / Unverified
    https://api.vsllm.com
  • Endpoint 3Historical / Unverified
    https://for.shuo.bar
Claim this provider

Verify ownership to unlock provider management features:

  • Edit provider name, content, and links
  • Get featured with priority traffic and visibility boost
  • Display a verified badge to build user trust

Leaderboard Rankings

Rankings are based on community-submitted tests and periodic health probes. Advisory only, not official data.

About VSLLM

VSLLM currently presents a New API-based gateway on vsllm.com. A live review on 2026-04-27 showed the default New API title and the standard /logo.png icon path, so LMSpeed keeps the provider copy conservative: VSLLM appears to operate a unified AI gateway for routing requests to multiple model families through one endpoint and management panel rather than the broader marketplace positioning previously listed.

Health Check

100%Recent availability
History (72 pts)
PastNow

API Benchmarks & Pricing

Compare 197 model rows across audit recency, latest speed tests, throughput, latency, and per-token pricing.

Model
Input ($/M)
Speed
Latency
doubao-seed-2-0-prodefault
$0.073/request--
$0.365/request38.6 t/s5.06 s
kimi-k3default
$0.146/request--
$0.029/request150.9 t/s3.07 s
$0.036/request--
$0.219/request43.7 t/s7.13 s
grok-4.5default
$0.073/request--
claude-fake-5default
$0.073/request--
$0.036/request--
glm-5.2-anthropicdefault
$0.073/request--
glm-5.2default
$0.073/request--
$0.015/request--
$0.730/request--
doubao-seed-code-preview-251028
---
gemini-3-pro-image-preview
---
gemini-3.1-flash-image-preview
---
glm-5
---
glm-5-turbo
---
glm-5.1
---
glm-5.1-free
---
gpt-image-2-codex
---
grok-4.20-0309-reasoning
---
hunyuan-2.0-thinking-20251109
---
Pro/deepseek-ai/DeepSeek-V3.2
---
Pro/zai-org/GLM-5.1
---

Showing 25 of 197 model rows

Recent Test Records

TimeModelSpeedLatency
Aug 12, 08:23 AM
gpt-5.6-luna
150.93 tok/s
3.07s
Jul 30, 01:08 PM
claude-opus-5
38.63 tok/s
5.06s
Jul 30, 01:05 PM
gpt-5.6-sol
43.72 tok/s
7.13s
Jun 10, 01:26 AM
claude-opus-4-7
32.27 tok/s
4.93s
May 30, 05:43 AM
claude-opus-4-8
28.20 tok/s
26.33s
May 30, 05:42 AM
claude-opus-4-8-free
44.93 tok/s
4.15s
May 21, 01:44 AM
gpt-5.5-pro20x
50.29 tok/s
1.52s
May 21, 01:42 AM
claude-opus-4-7-request
38.20 tok/s
1.84s
May 21, 01:40 AM
gemini-3.5-flash-antigravity
320.41 tok/s
7.77s
May 19, 08:11 AM
gemini-3-flash-preview-request-antigravity
193.61 tok/s
12.37s

Similar API Provider Alternatives to Compare

Compare VSLLM alternatives against 6 nearby API providers using 444 LMSpeed signals across shared model coverage, pricing, benchmark speed, uptime, and free-model availability.

ProviderWhy compareModelsFreeAvg priceSpeed30d uptime
VSLLM

vsllm-com

VSLLM runs a New API-powered AI gateway on vsllm.com for aggregated model access through a single endpoint.

Current provider baseline2092$0.056/M174 tok/s9960%
Dext API

ai-dext-top

Dext API is an OpenAI-compatible API relay offering access to 500+ LLM models at competitive prices, with multi-channel routing and broad model coverage.

  • Lower average pricing
  • More free-model options
  • Broader model coverage
  • Same provider category
667219$0.017/M19 tok/s3130%
钠 API

naapi-cc

Na API (naapi.cc) is an OpenAI-compatible LLM API gateway with competitive pricing and stable access to 100+ models from OpenAI, Anthropic, Google, and more.

  • Lower average pricing
  • Faster measured speed
  • Broader model coverage
  • Same provider category
5660$0.0031/M649 tok/s9960%
6345ywz API

api-6345ywz-cn

6345ywz API is an OpenAI-compatible API relay providing access to multiple AI models with competitive pricing.

  • Lower average pricing
  • Faster measured speed
  • Broader model coverage
  • Same provider category
4660$0.0000686/M272 tok/s9750%
RenRen API

llm-whitedream-top

RenRen API runs a New API-powered gateway on llm.whitedream.top for aggregated access to multiple AI models.

  • Lower average pricing
  • Faster measured speed
  • Broader model coverage
  • Same provider category
4590$0.0050/M271 tok/s9840%
Cuz AI

ai-cuz-lab-space

Cuz AI runs an OpenAI-compatible relay at ai.cuz-lab.space with broad model coverage, public pricing, and stable throughput for chat and coding workloads.

  • Lower average pricing
  • Higher 30-day availability
  • Broader model coverage
  • Same provider category
3930$0.020/M170 tok/s9980%
91VIP

91vip-futureppo-top

91VIP is a non-profit API service providing access to various AI models including Codex, Claude Code, and Open Code, with specific unlimited-use groups.

  • Lower average pricing
  • Faster measured speed
  • Broader model coverage
  • Same provider category
3832$0.054/M301 tok/s0%

Announcements

default8/31/2026

📢 The final phase of the underlying scheduling system has now been successfully completed.

This full upgrade has completely restructured the core routing strategy and high-concurrency processing mechanism. Currently, all system performance indicators have reached the expected optimal state, and the stability of the entire chain has been greatly improved. Thank you for your patience and understanding during the upgrade. Welcome to experience a more stable and efficient service ❤️

default8/29/2026

📢 The underlying scheduling system upgrade progress of this site has reached 2/3🎉. The load balancing and stability optimization of the core link have taken initial effect. The remaining difficult work is being urgently advanced. We will complete the full upgrade as soon as possible. Thank you for your patience and support.

default8/29/2026

📢 Another statement about the 'July 23 Poisoning Incident'

Dear friends, let me give you a detailed review of the 'poisoning incident' from a month ago! The incident occurred at 22:00 Beijing time on July 23. The hacker Gong Mouhua (extremely cunning, but we have fully obtained his personal mobile phone number, personal information, and his wife's contact information) cracked the backend password of our account pool and implanted malicious upstream. From 23:00 on July 23 to 8:00 on July 24, the poisoning took place. At midnight on the 24th, we discovered that the ccload password had been changed. Although we tried our best to intercept (the other party's upstream returned very quickly, and we intercepted part of it, but they kept changing tactics), the password was still changed at 3 a.m. Then we launched a massive cleanup💪🫵 and thoroughly reinforced and repaired ccload. It was not until 9:00 on July 24 that the issue was completely resolved.

It has been a month since this incident. Newcomers who don't know about it can check the announcement from that time. Of course, if that incident caused you any loss, please contact us immediately! We can directly report to the police and have this criminal arrested and locked up! 😡 We will absolutely fight to the end and never tolerate it. The system is now very safe, so feel free to use it.

default8/25/2026

📢 Folks, the peak-hour scheduling queue issue has theoretically been resolved! 🎉 However, since traffic hasn't peaked yet, we need to wait until tomorrow to further observe and verify the effect 🧐

default8/25/2026

📢 New Model Launch Notice

Folks, given the amazing performance of Gemini 3.7 Flash recently 🤩, we've added a super affordable version: gemini-3.7-flash-api!

Cost-effectiveness maxed out, experience just as smooth — go try it out 🚀✨

default8/25/2026

📢 Folks, we fixed a serious cascading bug 🛠️! Previously, Gemini's refusal/safety responses ("The prompt could not be submitted...") inherently lacked markers, so forwarding a refusal once would mistakenly ban and hard-isolate the account. The result: sending a prompt that gets refused would try and block all accounts in the pool one by one, causing requests to spin forever 😵‍💫. This issue is now fully fixed — the account pool won't be degraded by false positives anymore, use with confidence! 🚀

default8/25/2026

📢 Folks, we're urgently fixing an intermittent disconnection bug caused by high load 🛠️, please bear with us~ 🙏

default8/24/2026

Stable~

default8/23/2026

📢 Hey everyone, I'm here to strongly recommend our site's exclusive GPT channel! 🌟 Their GPT is super stable, daily calls are smooth as silk, and the best part is the price is ridiculously cheap at just 0.085x multiplier! 💰 If you want cheap and stable GPT, just go for it blindly. Portal: sub.unsee.you, hurry up and grab the deal! 🚀✨

default8/23/2026

📢 New Model Launch Notice

Hey everyone, given the amazing performance of Gemini 3.7 Flash recently 🤩, we've added a super affordable version: gemini-3.7-flash-api!

Cost-effectiveness is maxed out, experience is just as smooth, go try it out 🚀✨

default8/22/2026

📢 Plan Adjustment Announcement

Thank you all for your continued support and feedback.

After collecting and analyzing recent user feedback, we realized that the original plan pricing and quota allocation were unreasonable, leading to low cost-effectiveness and poor experience for some users. We sincerely apologize for this and have restructured and adjusted the entire plan system.

This adjustment adds the following three plans:

Plan NamePriceReset CycleSingle Quota
Basic Plan¥8/monthWeekly500
Light Plan¥12/monthEvery 5 hours200
Youth Plan¥32/monthEvery 5 hours500

The original "See You in 5 Hours" (¥49) single quota is the same as the newly launched "Youth Plan" (¥32), with a significantly lower price and greatly improved cost-effectiveness.

Original Plan Adjustments

Some original plans have significantly increased quotas while prices have been adjusted upward accordingly:

Adjustment ItemBefore AdjustmentAfter AdjustmentChange
Daily Card (Small Rush)¥6 / 500 quota¥9 / 1000 quotaPrice +50%, quota doubled

The core goal of this adjustment is: to make every penny count, to make the plan tiers more reasonable, and to allow users with different needs to find a suitable tier.

If you have any questions or suggestions about the plan adjustment, feel free to contact us. Thank you for your understanding and support 🙏

default8/21/2026

📢 Hey folks, we've quietly launched a super hot mysterious cutting-edge model stealth/ox-alpha 🥷! Its origins are extremely mysterious (the developer is completely anonymous), but it's incredibly powerful, with a 1M ultra-long context, making coding and running Agents silky smooth 🚀. It's fresh out of the oven, so go ahead and try it out, and guess whose alias it might be 🤫✨!

default8/20/2026

📢 Reminder for endpoint usage: Gemini must use the Gemini-specific endpoint, GPT series must use the /v1/responses endpoint. For other models (e.g., domestic models, Grok, etc.), just use the regular Chat endpoint. Please pay attention when configuring. 🛠️✨

default8/20/2026

📢 The issue with Grok's slow first token has been resolved! Response is now smooth 🚀✨

default8/18/2026

Hey folks, here's a progress update 🤗. GLM-5.3 hasn't been open-sourced yet, so deployment is slow and it's still unstable. We're currently fighting high load on the account pool scheduler. Actually, we have plenty of accounts, but since last Wednesday a scheduling bug caused the load to spike, triggering infinite system restarts. The worst part? It blew through 176TB of traffic! That huge traffic sink has been completely fixed now, but there are still some minor load issues to wrap up. Thanks for your patience 🥹❤️

default8/8/2026

📢 The new account pool pre-check system is currently undergoing high-concurrency stress testing. As a thank you for your patience, all subscription quotas have been reset as compensation! 👌 Thanks for your support!

default8/8/2026

📢 Endpoint reminder: Due to the scheduling system update, when calling the entire Grok model family and [opencode]deepseek-v4-flash, please prioritize using the /v1/responses endpoint; otherwise you'll be stuck waiting for the first token 😵‍💫! (The new scheduler is very fast – grok-4.5 first token in about 1.7s 🚀)

The endpoint limitation issue for the above models has been fully fixed – no need to switch endpoints anymore! Playground is also adapted.

Notes

  • Health checks: Scope: the 72-hour chart and recent availability measure API connectivity only. Each bar summarizes one hour of checks. Targets: LMSpeed tries the configured health check URL and provider status URL first, then API endpoints derived from known API hosts and recent speed-test base URLs. A website host is considered only when it looks like an API endpoint. Probe steps: each candidate goes through DNS lookup, TCP connection, TLS handshake for HTTPS, and an HTTP HEAD request with redirects followed. Probing stops after the first reachable candidate. Reachable criteria: every required network step must succeed. An HTTP response below 500 is treated as reachable, including 401 because it confirms that an authenticated API endpoint responded, except for statuses classified as blocked. Blocked results: HTTP 403, 429, 521, 525, and 530, plus detected WAF or Cloudflare challenges, are shown as blocked and excluded from availability calculations because LMSpeed cannot determine whether the API itself is down. Model availability: when a dedicated test key is configured, LMSpeed sends an authenticated GET request to a derived /models endpoint and compares returned model IDs with this provider's listed models. These per-model results appear in Models & Pricing and are not included in the provider connectivity percentage. Timeouts: TCP connection, TLS handshake, HTTP connectivity, and model requests each use a 20-second timeout. A full run can take longer when several candidates are tried. Frequency: a background worker checks all providers every 5 minutes by default. The 72-hour chart combines those samples into hourly bars, and the schedule may be changed by the service operator. Limit: automated samples are not an SLA and do not guarantee account quota, every model, every region, or successful completion requests. Check the provider's own status page before making operational decisions.
  • Domain Rating data is sourced from Ahrefs. It is a 0–100 backlink-based domain strength signal and does not measure API speed or reliability.
  • Announcements and FAQ are read from this provider's NewAPI status snapshot when available. LMSpeed stores the original content and optional English translations from the provider status source, then shows the localized fields on this page.