gemma4:31b
40.9
output tokens / second
125 tokens · 3.2s request
Output-token throughput benchmarks every 5 minutes
GLM-5.2 analysis of the rolling four-hour, five-minute dataset
Across the four-hour window, deepseek-v4-flash achieved the strongest average throughput at 103.84 token/s, while nemotron-3-ultra was the weakest at 45.14 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from a peak of 216.98 token/s down to 13.59 token/s at 04:30, reflecting a coefficient of variation of 53.4 percent. This dataset contains 288 valid observations across six models, corresponding to exactly 100.0 percent coverage of the 48 expected samples per model, meaning there are no missing-data limitations to report.