Real-world LLM speed leaderboard
Anonymous real-world speed results from TOKRACE, annotated by sample confidence. Current 9 models · 516 runs
ModelMedian tok/sTTFTPeak
Speed vs cost
Full pricing table →Median output speed against output price. Top-left is best (fast & cheap); the dashed line marks the Pareto frontier — models not beaten on both speed and price.
· Cost uses each model's output price per 1M tokens (CNY converted to USD); speed is the median from the leaderboard above.
Popular comparisons
Head-to-head speed pages for the top models on the board.
kimi-for-coding vs step-3.7-flashdeepseek-v4-flash vs kimi-for-codingkimi-for-coding vs mimo-v2.5deepseek-v4-pro vs kimi-for-codingkimi-for-coding vs MiniMax-M3glm-5.1 vs kimi-for-codingglm-5.2 vs kimi-for-codingdeepseek-v4-flash vs step-3.7-flashmimo-v2.5 vs step-3.7-flashdeepseek-v4-pro vs step-3.7-flashMiniMax-M3 vs step-3.7-flashglm-5.1 vs step-3.7-flashglm-5.2 vs step-3.7-flashdeepseek-v4-flash vs mimo-v2.5deepseek-v4-flash vs deepseek-v4-prodeepseek-v4-flash vs MiniMax-M3deepseek-v4-flash vs glm-5.1deepseek-v4-flash vs glm-5.2deepseek-v4-pro vs mimo-v2.5mimo-v2.5 vs MiniMax-M3glm-5.1 vs mimo-v2.5glm-5.2 vs mimo-v2.5deepseek-v4-pro vs MiniMax-M3deepseek-v4-pro vs glm-5.1deepseek-v4-pro vs glm-5.2glm-5.1 vs MiniMax-M3glm-5.2 vs MiniMax-M3glm-5.1 vs glm-5.2
· Data comes from voluntary anonymous sharing and contains speed metrics only, not prompts or API keys · Updates every 5 minutes
· Speed is affected by network, time of day and provider load; different endpoints for the same model are tracked separately