4-hour window · 96 points · 100.0% coverage
- gemma4:31b is the strongest model at 124.79 token/s average throughput; nemotron-3-ultra is the weakest at 37.77 token/s average, with a minimum of 3.19 token/s.
- glm-5.3 shows the steepest decline, trending -32.4% from 142.52 token/s at 21:20 to 92.24 token/s at 01:00, while nemotron-3-ultra is the most volatile (cv 81.9%, range 3.19–112.61 token/s); deepseek-v4-flash rose 27.5% to 120.54 token/s.
- No missing-data limitation: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%.