4-hour window · 384 points · 100.0% coverage
- glm-5.3 is the strongest model at 152.93 token/s average throughput (p95 180.97 token/s), while nemotron-3-ultra is the weakest at 5.45 token/s average, peaking at only 9.57 token/s.
- glm-5.2 shows the highest volatility (CV 34.8%), dropping to 7.02 token/s at 03:35 and 10.8 token/s at 04:20; glm-5.3 also dipped sharply to 36.92 token/s at 04:20. deepseek-v4-pro was the steadiest performer (CV 16.1%, stddev 16.85 token/s).
- No missing-data limitation applies: all eight models recorded 48 of 48 expected samples, 100.0% coverage, and 384 valid points across the four-hour window.