4-hour window · 84 points · 87.5% coverage
- gemma4:31b is the strongest model on average throughput at 101.61 token/s, while nemotron-3-ultra is the weakest at 22.13 token/s, roughly a fifth of the leader's pace.
- glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 76.1% and swings from 11.59 to 203.8 token/s within the window; glm-5.3 also declined 21.7% and nemotron-3-ultra fell 37.0% over the period.
- deepseek-v4-flash reported zero samples across all 12 expected observations, leaving 0% coverage and reducing overall dataset coverage to 87.5% (84 of 96 expected points), so its throughput is unassessed.