Analysis of the rolling four-hour, twenty-minute dataset
AI-generated · hourly
4-hour window · 96 points · 100.0% coverage
deepseek-v4.1-flash is the strongest model at 186.75 token/s average throughput (peak 235.93 token/s), while nemotron-3-ultra is the weakest at 22.48 token/s average (minimum 3.15 token/s).
glm-5.2 shows the most extreme volatility, ranging from 10.78 to 216.7 token/s with a 94.6% coefficient of variation; gemma4:31b also swung from 146.13 down to 43.95 token/s late in the window.
No missing-data limitation exists: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the four-hour window.