Real-world LLM speed leaderboard

Anonymous real-world speed results from TOKRACE, annotated by sample confidence. Current 11 models · 581 runs

ModelMedian tok/sTTFTPeak
api.kimi.com· 3 runsLow sample
224
avg 231
1.27s
467
api.kimi.com· 129 runsStable sample
184
avg 225
1.49s
900
api.stepfun.com· 69 runsStable sample
170
avg 178
1.99s
2,067
api.deepseek.com· 88 runsStable sample
142
avg 132
0.80s
638
api.minimaxi.com· 6 runsLow sample
100
avg 97
1.52s
395
api.xiaomimimo.com· 45 runsUsable signal
86
avg 88
4.38s
514
api.deepseek.com· 55 runsStable sample
85
avg 86
1.29s
396
open.bigmodel.cn· 52 runsStable sample
53
avg 61
4.64s
594
open.bigmodel.cn· 81 runsStable sample
51
avg 60
3.90s
846
api.xiaomimimo.com· 41 runsUsable signal
43
avg 48
1.98s
225
api.kimi.com· 12 runsUsable signal
36
avg 46
3.09s
157

Median output speed against output price. Top-left is best (fast & cheap); the dashed line marks the Pareto frontier — models not beaten on both speed and price.

062123185247$0.30$1$3$10Output $/1M (log, ← cheaper)Median tok/s (faster ↑)↖ fast & cheap = bestkimi-for-coding-highspeed · api.kimi.com speed 224 tok/s · output $4/1M · Pareto frontierkimi-for-coding-highspeedkimi-for-coding · api.kimi.com speed 184 tok/s · output $4/1Mstep-3.7-flash · api.stepfun.com speed 170 tok/s · output $1/1M · Pareto frontierstep-3.7-flashdeepseek-v4-flash · api.deepseek.com speed 142 tok/s · output $0.28/1M · Pareto frontierdeepseek-v4-flashMiniMax-M3 · api.minimaxi.com speed 100 tok/s · output $1/1Mmimo-v2.5 · api.xiaomimimo.com speed 86 tok/s · output $0.28/1M · Pareto frontiermimo-v2.5deepseek-v4-pro · api.deepseek.com speed 85 tok/s · output $0.87/1Mglm-5.1 · open.bigmodel.cn speed 53 tok/s · output $4/1Mglm-5.2 · open.bigmodel.cn speed 51 tok/s · output $4/1Mmimo-v2.5-pro · api.xiaomimimo.com speed 43 tok/s · output $0.28/1MKimi K3 · api.kimi.com speed 36 tok/s · output $14/1M

· Cost uses each model's output price per 1M tokens (CNY converted to USD); speed is the median from the leaderboard above.

Popular comparisons

Head-to-head speed pages for the top models on the board.

· Data comes from voluntary anonymous sharing and contains speed metrics only, not prompts or API keys · Updates every 5 minutes

· Speed is affected by network, time of day and provider load; different endpoints for the same model are tracked separately