Analysis of the rolling four-hour, twenty-minute dataset
AI-generated · hourly
4-hour window · 96 points · 100.0% coverage
deepseek-v4.1-flash is the strongest model at an average 184.46 token/s (peaking at 235.93 token/s), while nemotron-3-ultra is the weakest at an average 25.54 token/s, never exceeding 52.47 token/s.
glm-5.2 shows the most operationally significant volatility, with a 64.2% coefficient of variation and swings from 10.78 to 96.62 token/s within the four hours; nemotron-3-ultra is similarly erratic at 57.9%.
No missing-data limitation exists: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the period.