mimo-v2.5 vs qwen3.8-27b-heretic-ara-16gb-vram-xs-mtp speed comparison

Based on 59 anonymous user runs.

Verdict: mimo-v2.5 has faster output (median 84 vs 67 tok/s); mimo-v2.5 has faster TTFT (4.57s vs 7.68s).
Share and embed

Post to social channels, or use Markdown and badges for GitHub/README.

XFacebook微博LinkedIn
[![mimo-v2.5 is faster than qwen3.8-27b-heretic-ara-16gb-vram-xs-mtp: 84 tok/s on TOKRACE](https://www.tokrace.com/api/badge/compare/mimo-v2-5-vs-qwen3-8-27b-heretic-ara-16gb-vram-xs-mtp?locale=en)](https://www.tokrace.com/en/compare/mimo-v2-5-vs-qwen3-8-27b-heretic-ara-16gb-vram-xs-mtp)
Median output tok/s8467
Average output tok/s8567
TTFT4.57s7.68s
Peak tok/s514121
Samples572

· Data comes from voluntary anonymous sharing; medians reduce jitter · Updates every 5 minutes

· Speed is affected by network, time of day and provider load · Methodology

How to use this comparison

Writing/long output: Prioritize median output tok/s and peak speed.

Chat/agents: TTFT usually has a bigger UX impact.

Model selection: Rerun your real Prompt and inspect output quality too.

Run with current dataView full leaderboard

FAQ

Which model outputs faster, mimo-v2.5 or qwen3.8-27b-heretic-ara-16gb-vram-xs-mtp?

mimo-v2.5 has faster output (median 84 vs 67 tok/s); mimo-v2.5 has faster TTFT (4.57s vs 7.68s).

Why can output speed and TTFT have different winners?

Output tok/s measures sustained generation speed, while TTFT measures the wait until the first token. A model can generate long text faster while still taking longer to start.

How should I rerun this comparison?

Use the arena with the same Prompt, temperature and network conditions, then repeat a few times and combine the speed data with output quality.

Can I embed this comparison in GitHub or an article?

Yes. This page provides Markdown and HTML badges. The badge image URL is https://www.tokrace.com/api/badge/compare/mimo-v2-5-vs-qwen3-8-27b-heretic-ara-16gb-vram-xs-mtp?locale=en.