VSLLM currently presents a New API-based gateway on vsllm.com. A live review on 2026-04-27 showed the default New API title and the standard /logo.png icon path, so LMSpeed keeps the provider copy conservative: VSLLM appears to operate a unified AI gateway for routing requests to multiple model families through one endpoint and management panel rather than the broader marketplace positioning previously listed.
VSLLM runs a New API-powered AI gateway on vsllm.com for aggregated model access through a single endpoint.
- Models
- 209 models
- From
- $0.110/M
- Speed
- 174 tok/s
- Updated
- 7/2/2026
- Latency
- 0.00 s
- Created At
- 3/8/2026
- Recharge Rate
- ¥7.30 per $1 quota
Features
Login Methods
API Endpoints
- Endpoint 1Historical / Unverified
https://vsllm.com - Endpoint 2Historical / Unverified
https://api.vsllm.com - Endpoint 3Historical / Unverified
https://for.shuo.bar
Verify ownership to unlock provider management features:
- Edit provider name, content, and links
- Get featured with priority traffic and visibility boost
- Display a verified badge to build user trust
Leaderboard Rankings
Rankings are based on community-submitted tests and periodic health probes. Advisory only, not official data.
About VSLLM
Health Check
API Benchmarks & Pricing
Compare 197 model rows across audit recency, latest speed tests, throughput, latency, and per-token pricing.
doubao-seed-2-0-prodefault | $0.073/request | - | - | ||||||||||
claude-opus-5default | $0.365/request | 38.6 t/s | 5.06 s | ||||||||||
kimi-k3default | $0.146/request | - | - | ||||||||||
gpt-5.6-lunadefault | $0.029/request | 150.9 t/s | 3.07 s | ||||||||||
gpt-5.6-terradefault | $0.036/request | - | - | ||||||||||
gpt-5.6-soldefault | $0.219/request | 43.7 t/s | 7.13 s | ||||||||||
grok-4.5default | $0.073/request | - | - | ||||||||||
claude-fake-5default | $0.073/request | - | - | ||||||||||
claude-sonnet-5default | $0.036/request | - | - | ||||||||||
glm-5.2-anthropicdefault | $0.073/request | - | - | ||||||||||
glm-5.2default | $0.073/request | - | - | ||||||||||
glm-5.2-freedefault | $0.015/request | - | - | ||||||||||
claude-fable-5default | $0.730/request | - | - | ||||||||||
doubao-seed-code-preview-251028 | - | - | - | ||||||||||
gemini-3-pro-image-preview | - | - | - | ||||||||||
gemini-3.1-flash-image-preview | - | - | - | ||||||||||
glm-5 | - | - | - | ||||||||||
glm-5-turbo | - | - | - | ||||||||||
glm-5.1 | - | - | - | ||||||||||
glm-5.1-free | - | - | - | ||||||||||
gpt-image-2-codex | - | - | - | ||||||||||
grok-4.20-0309-reasoning | - | - | - | ||||||||||
hunyuan-2.0-thinking-20251109 | - | - | - | ||||||||||
Pro/deepseek-ai/DeepSeek-V3.2 | - | - | - | ||||||||||
Pro/zai-org/GLM-5.1 | - | - | - |
Showing 25 of 197 model rows
Recent Test Records
| Time | Model | Speed | Latency |
|---|---|---|---|
| Aug 12, 08:23 AM | gpt-5.6-luna | 150.93 tok/s | 3.07s |
| Jul 30, 01:08 PM | claude-opus-5 | 38.63 tok/s | 5.06s |
| Jul 30, 01:05 PM | gpt-5.6-sol | 43.72 tok/s | 7.13s |
| Jun 10, 01:26 AM | claude-opus-4-7 | 32.27 tok/s | 4.93s |
| May 30, 05:43 AM | claude-opus-4-8 | 28.20 tok/s | 26.33s |
| May 30, 05:42 AM | claude-opus-4-8-free | 44.93 tok/s | 4.15s |
| May 21, 01:44 AM | gpt-5.5-pro20x | 50.29 tok/s | 1.52s |
| May 21, 01:42 AM | claude-opus-4-7-request | 38.20 tok/s | 1.84s |
| May 21, 01:40 AM | gemini-3.5-flash-antigravity | 320.41 tok/s | 7.77s |
| May 19, 08:11 AM | gemini-3-flash-preview-request-antigravity | 193.61 tok/s | 12.37s |
Similar API Provider Alternatives to Compare
Compare VSLLM alternatives against 6 nearby API providers using 444 LMSpeed signals across shared model coverage, pricing, benchmark speed, uptime, and free-model availability.
| Provider | Why compare | Models | Free | Avg price | Speed | 30d uptime |
|---|---|---|---|---|---|---|
| VSLLM vsllm-com VSLLM runs a New API-powered AI gateway on vsllm.com for aggregated model access through a single endpoint. | Current provider baseline | 209 | 2 | $0.056/M | 174 tok/s | 9960% |
| Dext API ai-dext-top Dext API is an OpenAI-compatible API relay offering access to 500+ LLM models at competitive prices, with multi-channel routing and broad model coverage. |
| 667 | 219 | $0.017/M | 19 tok/s | 3130% |
| 钠 API naapi-cc Na API (naapi.cc) is an OpenAI-compatible LLM API gateway with competitive pricing and stable access to 100+ models from OpenAI, Anthropic, Google, and more. |
| 566 | 0 | $0.0031/M | 649 tok/s | 9960% |
| 6345ywz API api-6345ywz-cn 6345ywz API is an OpenAI-compatible API relay providing access to multiple AI models with competitive pricing. |
| 466 | 0 | $0.0000686/M | 272 tok/s | 9750% |
| RenRen API llm-whitedream-top RenRen API runs a New API-powered gateway on llm.whitedream.top for aggregated access to multiple AI models. |
| 459 | 0 | $0.0050/M | 271 tok/s | 9840% |
| Cuz AI ai-cuz-lab-space Cuz AI runs an OpenAI-compatible relay at ai.cuz-lab.space with broad model coverage, public pricing, and stable throughput for chat and coding workloads. |
| 393 | 0 | $0.020/M | 170 tok/s | 9980% |
| 91VIP 91vip-futureppo-top 91VIP is a non-profit API service providing access to various AI models including Codex, Claude Code, and Open Code, with specific unlimited-use groups. |
| 383 | 2 | $0.054/M | 301 tok/s | 0% |
Announcements
📢 The final phase of the underlying scheduling system has now been successfully completed.
This full upgrade has completely restructured the core routing strategy and high-concurrency processing mechanism. Currently, all system performance indicators have reached the expected optimal state, and the stability of the entire chain has been greatly improved. Thank you for your patience and understanding during the upgrade. Welcome to experience a more stable and efficient service ❤️
📢 The underlying scheduling system upgrade progress of this site has reached 2/3🎉. The load balancing and stability optimization of the core link have taken initial effect. The remaining difficult work is being urgently advanced. We will complete the full upgrade as soon as possible. Thank you for your patience and support.
📢 Another statement about the 'July 23 Poisoning Incident'
Dear friends, let me give you a detailed review of the 'poisoning incident' from a month ago! The incident occurred at 22:00 Beijing time on July 23. The hacker Gong Mouhua (extremely cunning, but we have fully obtained his personal mobile phone number, personal information, and his wife's contact information) cracked the backend password of our account pool and implanted malicious upstream. From 23:00 on July 23 to 8:00 on July 24, the poisoning took place. At midnight on the 24th, we discovered that the ccload password had been changed. Although we tried our best to intercept (the other party's upstream returned very quickly, and we intercepted part of it, but they kept changing tactics), the password was still changed at 3 a.m. Then we launched a massive cleanup💪🫵 and thoroughly reinforced and repaired ccload. It was not until 9:00 on July 24 that the issue was completely resolved.
It has been a month since this incident. Newcomers who don't know about it can check the announcement from that time. Of course, if that incident caused you any loss, please contact us immediately! We can directly report to the police and have this criminal arrested and locked up! 😡 We will absolutely fight to the end and never tolerate it. The system is now very safe, so feel free to use it.
📢 Folks, the peak-hour scheduling queue issue has theoretically been resolved! 🎉 However, since traffic hasn't peaked yet, we need to wait until tomorrow to further observe and verify the effect 🧐
📢 New Model Launch Notice
Folks, given the amazing performance of Gemini 3.7 Flash recently 🤩, we've added a super affordable version: gemini-3.7-flash-api!
Cost-effectiveness maxed out, experience just as smooth — go try it out 🚀✨
📢 Folks, we fixed a serious cascading bug 🛠️! Previously, Gemini's refusal/safety responses ("The prompt could not be submitted...") inherently lacked markers, so forwarding a refusal once would mistakenly ban and hard-isolate the account. The result: sending a prompt that gets refused would try and block all accounts in the pool one by one, causing requests to spin forever 😵💫. This issue is now fully fixed — the account pool won't be degraded by false positives anymore, use with confidence! 🚀
📢 Folks, we're urgently fixing an intermittent disconnection bug caused by high load 🛠️, please bear with us~ 🙏
Stable~
📢 Hey everyone, I'm here to strongly recommend our site's exclusive GPT channel! 🌟 Their GPT is super stable, daily calls are smooth as silk, and the best part is the price is ridiculously cheap at just 0.085x multiplier! 💰 If you want cheap and stable GPT, just go for it blindly. Portal: sub.unsee.you, hurry up and grab the deal! 🚀✨
📢 New Model Launch Notice
Hey everyone, given the amazing performance of Gemini 3.7 Flash recently 🤩, we've added a super affordable version: gemini-3.7-flash-api!
Cost-effectiveness is maxed out, experience is just as smooth, go try it out 🚀✨
📢 Plan Adjustment Announcement
Thank you all for your continued support and feedback.
After collecting and analyzing recent user feedback, we realized that the original plan pricing and quota allocation were unreasonable, leading to low cost-effectiveness and poor experience for some users. We sincerely apologize for this and have restructured and adjusted the entire plan system.
This adjustment adds the following three plans:
| Plan Name | Price | Reset Cycle | Single Quota |
|---|---|---|---|
| Basic Plan | ¥8/month | Weekly | 500 |
| Light Plan | ¥12/month | Every 5 hours | 200 |
| Youth Plan | ¥32/month | Every 5 hours | 500 |
The original "See You in 5 Hours" (¥49) single quota is the same as the newly launched "Youth Plan" (¥32), with a significantly lower price and greatly improved cost-effectiveness.
Original Plan Adjustments
Some original plans have significantly increased quotas while prices have been adjusted upward accordingly:
| Adjustment Item | Before Adjustment | After Adjustment | Change |
|---|---|---|---|
| Daily Card (Small Rush) | ¥6 / 500 quota | ¥9 / 1000 quota | Price +50%, quota doubled |
The core goal of this adjustment is: to make every penny count, to make the plan tiers more reasonable, and to allow users with different needs to find a suitable tier.
If you have any questions or suggestions about the plan adjustment, feel free to contact us. Thank you for your understanding and support 🙏
📢 Hey folks, we've quietly launched a super hot mysterious cutting-edge model stealth/ox-alpha 🥷! Its origins are extremely mysterious (the developer is completely anonymous), but it's incredibly powerful, with a 1M ultra-long context, making coding and running Agents silky smooth 🚀. It's fresh out of the oven, so go ahead and try it out, and guess whose alias it might be 🤫✨!
📢 Reminder for endpoint usage: Gemini must use the Gemini-specific endpoint, GPT series must use the /v1/responses endpoint. For other models (e.g., domestic models, Grok, etc.), just use the regular Chat endpoint. Please pay attention when configuring. 🛠️✨
📢 The issue with Grok's slow first token has been resolved! Response is now smooth 🚀✨
Hey folks, here's a progress update 🤗. GLM-5.3 hasn't been open-sourced yet, so deployment is slow and it's still unstable. We're currently fighting high load on the account pool scheduler. Actually, we have plenty of accounts, but since last Wednesday a scheduling bug caused the load to spike, triggering infinite system restarts. The worst part? It blew through 176TB of traffic! That huge traffic sink has been completely fixed now, but there are still some minor load issues to wrap up. Thanks for your patience 🥹❤️
📢 The new account pool pre-check system is currently undergoing high-concurrency stress testing. As a thank you for your patience, all subscription quotas have been reset as compensation! 👌 Thanks for your support!
📢 Endpoint reminder: Due to the scheduling system update, when calling the entire Grok model family and [opencode]deepseek-v4-flash, please prioritize using the /v1/responses endpoint; otherwise you'll be stuck waiting for the first token 😵💫! (The new scheduler is very fast – grok-4.5 first token in about 1.7s 🚀)
The endpoint limitation issue for the above models has been fully fixed – no need to switch endpoints anymore! Playground is also adapted.
Notes
- Health checks: Scope: the 72-hour chart and recent availability measure API connectivity only. Each bar summarizes one hour of checks. Targets: LMSpeed tries the configured health check URL and provider status URL first, then API endpoints derived from known API hosts and recent speed-test base URLs. A website host is considered only when it looks like an API endpoint. Probe steps: each candidate goes through DNS lookup, TCP connection, TLS handshake for HTTPS, and an HTTP HEAD request with redirects followed. Probing stops after the first reachable candidate. Reachable criteria: every required network step must succeed. An HTTP response below 500 is treated as reachable, including 401 because it confirms that an authenticated API endpoint responded, except for statuses classified as blocked. Blocked results: HTTP 403, 429, 521, 525, and 530, plus detected WAF or Cloudflare challenges, are shown as blocked and excluded from availability calculations because LMSpeed cannot determine whether the API itself is down. Model availability: when a dedicated test key is configured, LMSpeed sends an authenticated GET request to a derived /models endpoint and compares returned model IDs with this provider's listed models. These per-model results appear in Models & Pricing and are not included in the provider connectivity percentage. Timeouts: TCP connection, TLS handshake, HTTP connectivity, and model requests each use a 20-second timeout. A full run can take longer when several candidates are tried. Frequency: a background worker checks all providers every 5 minutes by default. The 72-hour chart combines those samples into hourly bars, and the schedule may be changed by the service operator. Limit: automated samples are not an SLA and do not guarantee account quota, every model, every region, or successful completion requests. Check the provider's own status page before making operational decisions.
- Domain Rating data is sourced from Ahrefs. It is a 0–100 backlink-based domain strength signal and does not measure API speed or reliability.
- Announcements and FAQ are read from this provider's NewAPI status snapshot when available. LMSpeed stores the original content and optional English translations from the provider status source, then shows the localized fields on this page.

