Review Agent, Coding, and Reasoning model scores, compare provider rates, free tiers, and model coverage, measure API latency, throughput, and duration, and detect model, prompt, and error leakage risks.
Start with model score lookup, then compare per-token pricing, five-round speed tests, health checks, and safety probes before an API reaches production.
Search a model name from the homepage and jump into benchmark pages with AA score, coding, math, and model metadata.
Compare per-token pricing across 100+ providers, find cheaper APIs for each model, and track free tiers and credits.
Run a five-round benchmark with standardized prompts to measure first-token latency, output throughput, and response time.
Audit any OpenAI-compatible API for model authenticity, hidden prompts, instruction tampering, stream integrity, and error leakage, then share a plain-language report.
Enter a base URL, API key, and model ID to test official providers, proxies, relays, or self-hosted endpoints.
Review first-token latency, output throughput, total duration, health, and recent probe signals to judge stability.
How LMSpeed handles model benchmarks, provider pricing, speed tests, and API audits.
Use the model search on the homepage or open the Model Performance leaderboard. Search by model name to find the model detail page, where LMSpeed connects benchmark signals such as AA score, coding, and math with model metadata.
LMSpeed aggregates per-token pricing from 100+ API providers. Visit any model page to see a side-by-side pricing comparison table showing input and output rates per million tokens, so you can find the cheapest provider for each model.
Many providers offer free API tiers or credits for popular models like DeepSeek, Gemini, and Llama. Check our Free LLM API directory for a complete list of models with free access, including speed benchmarks for each free provider.
LMSpeed employs a five-round continuous stress testing mechanism with standardized prompts. Token calculations are performed accurately using tiktoken, measuring output throughput (tokens per second) and first-token latency.
It sends multiple safety probes to an OpenAI-compatible endpoint to check whether the model identity matches, hidden system prompts are injected, user instructions are rewritten, streaming responses stay intact, and errors leak sensitive implementation details. The result includes a risk score and a shareable report. API keys are only used for that audit and are not written to public reports.
Use our performance leaderboards and model detail pages to visually compare API speed benchmarks across providers. The system ranks providers by throughput, latency, and health, helping you choose the fastest and most reliable API.
Provider pages and the health leaderboard already show recent health checks, probe latency, success or failure status, and stability rankings. Broader continuous monitoring and alerting will keep expanding.
The newest relay audit reports where endpoint profile, model identity, prompt safety, and response integrity all scored 100.
A live cut of newly tracked models and benchmark leaders, focused on Artificial Analysis scores for overall intelligence, coding, and math.
| Model | Context | Input | Output | Providers | Agents | Coding | Reasoning | Knowledge | Math | Multilingual | Multimodal | Instruction following | Throughput | Latency | Release date |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 5AnthropicNEW | Context1M | Input$5.00/M | Output$25.00/M | Providers +74 | 67.5±8.2 | 70.1±6.6 | 60.6±8.7 | 67±14.0P | — | 64.3±17.1P | 58.5±17.0P | — | Throughput 39 t/s | Latency 5.06s | Release date2026-07-24 |
| Claude Fable 5Anthropic | Context1M | Input$10.00/M | Output$50.00/M | Providers +115 | 65.7±8.2 | 68.1±6.9 | 60.7±10.8E | 61.1±14.0P | — | — | 46.2±17.0P | 50.4±16.0P | Throughput 56 t/s | Latency 3.67s | Release date2026-06-09 |
| Kimi K3MoonshotAINEW | Context1.0M | Input$3.00/M | Output$15.00/M | Providers +70 | 68.9±7.3 | 59.6±11.9E | 58.8±10.8E | 61.9±14.0P | — | — | 67.9±11.3E | — | Throughput — | Latency — | Release date2026-07-16 |
| Qwen3.8 MaxQwenNEW | Context1M | Input$2.00/M | Output$6.00/M | Providers +12 | 65.8±11.3E | 67.3±11.2E | 65.9±11.3E | 51.6±16.1P | — | — | 65.9±6.3 | 57.8±16.0P | Throughput — | Latency — | Release date2026-08-03 |
| GPT-5.6 SolOpenAINEW | Context1.1M | Input$5.00/M | Output$30.00/M | Providers +128 | 68.2±7.3 | 63.2±9.3E | 61.7±8.7 | 57.2±16.7P | 67±16.1P | — | 59.9±16.1P | 53.5±16.0P | Throughput 44 t/s | Latency 4.27s | Release date2026-07-09 |
| Claude Opus 4.8Anthropic | Context1M | Input$5.00/M | Output$25.00/M | Providers +148 | 63.2±5.3 | 66±6.6 | 56.7±8.7 | 64.4±14.0P | 58±16.1P | 55±17.1P | 61.6±12.0E | 50±16.0P | Throughput 232 t/s | Latency 2.21s | Release date2026-05-27 |
| Muse Spark 1.2MetaNEW | Context1.0M | Input$1.25/M | Output$4.25/M | Providers | 62.3±16.0P | 64.6±16.0P | 61.4±13.9P | — | — | — | — | — | Throughput — | Latency — | Release date2026-08-05 |
| GPT-5.5OpenAI | Context1.1M | Input$5.00/M | Output$30.00/M | Providers +143 | 62.7±5.1 | 60±8.7 | 59.6±8.7 | 60.4±14.0P | 56.9±16.1P | — | 55.9±16.1P | 54.7±16.0P | Throughput 46 t/s | Latency 5.05s | Release date2026-04-24 |
| Grok 4.5SpaceXAI | Context500K | Input$2.00/M | Output$6.00/M | Providers +100 | 61.9±8.4 | 59±6.9 | 54.5±8.7 | 58.1±14.0P | — | — | 48.5±17.0P | — | Throughput 55 t/s | Latency 8.64s | Release date2026-07-08 |
| Claude Sonnet 5Anthropic | Context1M | Input$2.00/M | Output$10.00/M | Providers +105 | 59.8±8.2 | 59.1±6.6 | 56.5±10.8E | 61.7±14.0P | — | — | 60.3±16.2P | — | Throughput — | Latency — | Release date2026-06-30 |
| Claude Opus 4.7Anthropic | Context1M | Input$5.00/M | Output$25.00/M | Providers +192 | 54.9±6.9 | 60.4±11.2E | 56.3±10.8E | 54.3±14.0P | 56.8±16.1P | — | — | 44.5±16.0P | Throughput 47 t/s | Latency 4.89s | Release date2026-05-12 |
| Muse Spark 1.1MetaNEW | Context1.0M | Input$1.25/M | Output$4.25/M | Providers | 60.4±7.1 | 58.1±9.3E | 61.4±10.2E | 64.7±14.0P | — | — | 60.5±16.2P | — | Throughput — | Latency — | Release date2026-07-16 |
| GLM-5.2Z.ai | Context1.0M | Input$1.30/M | Output$4.18/M | Providers +136 | 60.5±7.2 | 54.4±8.9 | 54.9±10.8E | 58.7±14.0P | 70.8±11.5E | — | — | 53.7±16.0P | Throughput 61 t/s | Latency 7.79s | Release date2026-06-16 |
| GPT-5.6 LunaOpenAINEW | Context1.1M | Input$0.200/M | Output$1.20/M | Providers +121 | 59.8±7.4 | 56.9±9.3E | 54.9±8.7 | 56.5±16.7P | 62.4±16.1P | — | 50.2±16.1P | — | Throughput 22 t/s | Latency 18.05s | Release date2026-07-09 |
| Gemini 3.5 FlashGoogle | Context1.0M | Input$1.50/M | Output$9.00/M | Providers +102 | 61.1±5.6 | 56.1±8.9 | 53.5±8.4 | 55.4±14.0P | 55.1±16.1P | — | 57.3±11.9E | 54.9±16.0P | Throughput 425 t/s | Latency 4.10s | Release date2026-05-19 |
| Gemini 3.6 FlashGoogleNEW | Context1.0M | Input$1.50/M | Output$7.50/M | Providers +59 | 53±9.0 | 54.7±11.9E | 59±10.8E | 58±14.0P | — | — | 51.3±17.0P | — | Throughput — | Latency — | Release date2026-07-21 |
Compare API pricing, speed benchmarks, and performance data across providers.
A unified API gateway providing access to multiple large language models with direct connectivity in China.
Health
100%
Tests
145
Last check
Aug 8
API price
No health checks yet
Health
100%
Tests
5
Last check
Aug 8
API price
No health checks yet
Health
100%
Tests
5
Last check
Aug 8
API price
No health checks yet
Health
99%
Tests
15
Last check
Aug 8
API price
No health checks yet
Ollama provides a platform to run and integrate open-source AI models locally or in the cloud.
Health
100%
Tests
105
Last check
Aug 8
API price
No health checks yet
DeepSeek provides API access to its latest large language models for text generation and coding tasks.
Health
100%
Tests
702
Last check
Aug 8
API price
No health checks yet
OpenCode is an open-source AI coding agent that integrates with terminals, IDEs, and desktop apps, supporting multiple models and providers.
Health
100%
Tests
225
Last check
Aug 8
API price
No health checks yet
Provides cost-effective generative AI cloud services based on open-source models for text, image, video, and audio generation.
Health
73%
Tests
871
Last check
Aug 8
API price
No health checks yet
Zhipu AI provides the GLM series of large language models including GLM-4, ChatGLM, and CodeGeeX for text, code, and multimodal tasks.
Health
100%
Tests
400
Last check
Aug 8
API price
No health checks yet
Health
100%
Tests
20
Last check
Aug 8
API price
No health checks yet
Cuz AI runs an OpenAI-compatible relay at ai.cuz-lab.space with broad model coverage, public pricing, and stable throughput for chat and coding workloads.
XiaMiAPI is a unified LLM API gateway offering access to multiple AI models with competitive pricing and stable performance.
Health
86%
Tests
10
Last check
Aug 8
API price
No health checks yet