Real-world LLM speed leaderboard
Anonymous real-world speed results from TOKRACE, annotated by sample confidence. Current 21 models · 719 runs
ModelMedian tok/sTTFTPeak
management-cheats-steve-mercy.trycloudflare.com· 2 runsLow sample
67
avg 67
7.68s
121
Speed vs cost
Full pricing table →Median output speed against output price. Top-left is best (fast & cheap); the dashed line marks the Pareto frontier — models not beaten on both speed and price.
· Cost uses each model's output price per 1M tokens (CNY converted to USD); speed is the median from the leaderboard above.
Popular comparisons
Head-to-head speed pages for the top models on the board.
deepseek/deepseek-v4-flash-free vs kimi-for-coding-highspeedkimi-for-coding vs kimi-for-coding-highspeedkimi-for-coding-highspeed vs step-3.7-flashDeepSeek V4.1 Flash vs kimi-for-coding-highspeedkimi-for-coding-highspeed vs z-ai/glm-5.3-flash-freedeepseek-v4-flash vs kimi-for-coding-highspeedagnes-2.5-flash vs kimi-for-coding-highspeeddeepseek/deepseek-v4-flash-free vs kimi-for-codingdeepseek/deepseek-v4-flash-free vs step-3.7-flashdeepseek/deepseek-v4-flash-free vs DeepSeek V4.1 Flashdeepseek/deepseek-v4-flash-free vs z-ai/glm-5.3-flash-freedeepseek/deepseek-v4-flash-free vs deepseek-v4-flashagnes-2.5-flash vs deepseek/deepseek-v4-flash-freekimi-for-coding vs step-3.7-flashDeepSeek V4.1 Flash vs kimi-for-codingkimi-for-coding vs z-ai/glm-5.3-flash-freedeepseek-v4-flash vs kimi-for-codingagnes-2.5-flash vs kimi-for-codingDeepSeek V4.1 Flash vs step-3.7-flashstep-3.7-flash vs z-ai/glm-5.3-flash-freedeepseek-v4-flash vs step-3.7-flashagnes-2.5-flash vs step-3.7-flashDeepSeek V4.1 Flash vs z-ai/glm-5.3-flash-freeDeepSeek V4.1 Flash vs deepseek-v4-flashagnes-2.5-flash vs DeepSeek V4.1 Flashdeepseek-v4-flash vs z-ai/glm-5.3-flash-freeagnes-2.5-flash vs z-ai/glm-5.3-flash-freeagnes-2.5-flash vs deepseek-v4-flash
· Data comes from voluntary anonymous sharing and contains speed metrics only, not prompts or API keys · Updates every 5 minutes
· Speed is affected by network, time of day and provider load; different endpoints for the same model are tracked separately