4-hour window · 96 points · 100.0% coverage
- deepseek-v4.1-flash is the strongest model at 146.9 token/s average throughput, peaking at 237.72 token/s; nemotron-3-ultra is the weakest at 24.72 token/s average, never exceeding 40.01 token/s.
- glm-5.2 shows the most operationally significant volatility, with a 79.2% coefficient of variation, a single 203.28 token/s spike at 19:00 against a 61.38 token/s average, and a -21.8% trend; deepseek-v4-pro is steadiest at 13.7% CV.
- No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.