4-hour window · 96 points · 100.0% coverage
- deepseek-v4-flash is the strongest model at 111.82 token/s average throughput, narrowly ahead of gemma4:31b at 110.42 token/s and deepseek-v4-pro at 108.16 token/s; nemotron-3-ultra is the weakest at 3.71 token/s average, never exceeding 6.96 token/s.
- glm-5.3 shows the highest volatility (cv 45.7%), ranging from 64.79 to 189.77 token/s, and glm-5.3-flash dropped to 13.93 token/s at 04:00 before rebounding to 129.99 token/s at 04:20; deepseek-v4-flash trended up 36.4% to a 150.70 token/s peak.
- No samples are missing: all eight models report 12 of 12 expected observations, 96 valid points, and 100.0% coverage, so the only limitation is the short four-hour window.