gemma4:31b
62.7
output tokens / second
123 tokens · 2.1s request
Output-token throughput benchmarks every 5 minutes
GLM-5.2 analysis of the rolling four-hour, five-minute dataset
Across the four-hour window, glm-5.2 had the highest average output-token throughput at 122.38 token/s, while nemotron-3-ultra was weakest at 46.77 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from a peak of 223.93 token/s at 02:25 to a low of 13.59 token/s at 04:30, reflecting a 60.7 percent downward trend. In contrast, deepseek-v4-flash maintained the most stable performance, averaging 106.89 token/s with a coefficient of variation of 20.7 percent. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.