Real-world LLM speed leaderboard

Anonymous real-world speed results from TOKRACE, annotated by sample confidence. Current 21 models · 719 runs

ModelMedian tok/sTTFTPeak
api.kimi.com· 6 runsLow sample
250
avg 257
1.31s
467
api.orcarouter.ai· 2 runsLow sample
175
avg 175
1.70s
1,688
api.kimi.com· 136 runsStable sample
171
avg 224
1.73s
900
api.stepfun.com· 69 runsStable sample
170
avg 178
1.99s
2,067
api.deepseek.com· 2 runsLow sample
152
avg 152
0.34s
191
api.orcarouter.ai· 2 runsLow sample
142
avg 142
1.51s
856
api.deepseek.com· 113 runsStable sample
139
avg 135
1.08s
638
api.agnes-ai.cn· 11 runsUsable signal
125
avg 106
1.13s
351
api.orcarouter.ai· 1 runsLow sample
121
avg 121
1.28s
150
api.deepseek.com· 1 runsLow sample
105
avg 105
0.57s
113
open.bigmodel.cn· 2 runsLow sample
105
avg 105
1.15s
1,165
api.orcarouter.ai· 3 runsLow sample
101
avg 144
1.53s
1,294
api.minimaxi.com· 6 runsLow sample
100
avg 97
1.52s
395
api.deepseek.com· 75 runsStable sample
85
avg 88
1.65s
396
api.xiaomimimo.com· 59 runsStable sample
83
avg 85
4.53s
514
open.bigmodel.cn· 3 runsLow sample
73
avg 72
0.96s
477
management-cheats-steve-mercy.trycloudflare.com· 2 runsLow sample
67
avg 67
7.68s
121
open.bigmodel.cn· 61 runsStable sample
59
avg 65
4.51s
594
open.bigmodel.cn· 91 runsStable sample
55
avg 64
3.93s
846
api.xiaomimimo.com· 55 runsStable sample
44
avg 50
2.73s
225
api.kimi.com· 19 runsUsable signal
36
avg 43
3.33s
157

Median output speed against output price. Top-left is best (fast & cheap); the dashed line marks the Pareto frontier — models not beaten on both speed and price.

069138206275$0.30$1$3$10Output $/1M (log, ← cheaper)Median tok/s (faster ↑)↖ fast & cheap = bestkimi-for-coding-highspeed · api.kimi.com speed 250 tok/s · output $4/1M · Pareto frontierkimi-for-coding-highspeeddeepseek/deepseek-v4-flash-free · api.orcarouter.ai speed 175 tok/s · output $0.28/1M · Pareto frontierdeepseek/deepseek-v4-flash-freekimi-for-coding · api.kimi.com speed 171 tok/s · output $4/1Mstep-3.7-flash · api.stepfun.com speed 170 tok/s · output $1/1Mdeepseek-v4-flash · api.deepseek.com speed 139 tok/s · output $0.28/1Mdeepseek-chat · api.deepseek.com speed 105 tok/s · output $0.28/1MMiniMax-M3 · api.minimaxi.com speed 100 tok/s · output $1/1Mdeepseek-v4-pro · api.deepseek.com speed 85 tok/s · output $0.87/1Mmimo-v2.5 · api.xiaomimimo.com speed 83 tok/s · output $0.28/1M · Pareto frontiermimo-v2.5glm-5.1 · open.bigmodel.cn speed 59 tok/s · output $4/1Mglm-5.2 · open.bigmodel.cn speed 55 tok/s · output $4/1Mmimo-v2.5-pro · api.xiaomimimo.com speed 44 tok/s · output $0.28/1MKimi K3 · api.kimi.com speed 36 tok/s · output $14/1M

· Cost uses each model's output price per 1M tokens (CNY converted to USD); speed is the median from the leaderboard above.

Popular comparisons

Head-to-head speed pages for the top models on the board.

· Data comes from voluntary anonymous sharing and contains speed metrics only, not prompts or API keys · Updates every 5 minutes

· Speed is affected by network, time of day and provider load; different endpoints for the same model are tracked separately