4-hour window · 96 points · 100.0% coverage
- deepseek-v4.1-flash is the strongest model at 183.38 token/s average throughput (peak 227.33 token/s), while nemotron-3-ultra is the weakest at 21.14 token/s average, roughly one-ninth of the leader.
- glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 73.7 percent, swinging from 19.99 token/s at 21:40 to 191.19 token/s at 22:20; gemma4:31b also dipped to 50.18 token/s around 20:40.
- No missing-data limitation applies: all eight models have 12 of 12 samples and 100.0 percent coverage, with 96 valid points overall, so the four-hour window is fully represented.