Real-world LLM speed leaderboard

Anonymous real-world speed results from TOKRACE, annotated by sample confidence. Current 9 models · 516 runs

ModelMedian tok/sTTFTPeak
api.kimi.com· 119 runsStable sample
188
avg 235
1.48s
900
api.stepfun.com· 69 runsStable sample
170
avg 178
1.99s
2,067
api.deepseek.com· 76 runsStable sample
142
avg 133
0.79s
346
api.xiaomimimo.com· 39 runsUsable signal
86
avg 89
4.71s
514
api.deepseek.com· 48 runsUsable signal
80
avg 85
1.39s
293
api.minimaxi.com· 3 runsLow sample
80
avg 86
1.40s
179
open.bigmodel.cn· 51 runsStable sample
53
avg 61
4.70s
594
open.bigmodel.cn· 76 runsStable sample
48
avg 58
3.98s
475
api.xiaomimimo.com· 35 runsUsable signal
40
avg 43
2.03s
182

Median output speed against output price. Top-left is best (fast & cheap); the dashed line marks the Pareto frontier — models not beaten on both speed and price.

052103155207$0.30$1$3Output $/1M (log, ← cheaper)Median tok/s (faster ↑)↖ fast & cheap = bestkimi-for-coding · api.kimi.com speed 188 tok/s · output $4/1M · Pareto frontierkimi-for-codingstep-3.7-flash · api.stepfun.com speed 170 tok/s · output $1/1M · Pareto frontierstep-3.7-flashdeepseek-v4-flash · api.deepseek.com speed 142 tok/s · output $0.28/1M · Pareto frontierdeepseek-v4-flashmimo-v2.5 · api.xiaomimimo.com speed 86 tok/s · output $0.28/1M · Pareto frontiermimo-v2.5deepseek-v4-pro · api.deepseek.com speed 80 tok/s · output $0.87/1MMiniMax-M3 · api.minimaxi.com speed 80 tok/s · output $1/1Mglm-5.1 · open.bigmodel.cn speed 53 tok/s · output $4/1Mglm-5.2 · open.bigmodel.cn speed 48 tok/s · output $4/1Mmimo-v2.5-pro · api.xiaomimimo.com speed 40 tok/s · output $0.28/1M

· Cost uses each model's output price per 1M tokens (CNY converted to USD); speed is the median from the leaderboard above.

Popular comparisons

Head-to-head speed pages for the top models on the board.

· Data comes from voluntary anonymous sharing and contains speed metrics only, not prompts or API keys · Updates every 5 minutes

· Speed is affected by network, time of day and provider load; different endpoints for the same model are tracked separately