← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1209 retained summaries
4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 185.62 token/s, while nemotron-3-ultra was weakest at 35.94 token/s. The most operationally significant volatility is deepseek-v4-flash, which exhibited extreme swings between 19.96 and 155.36 token/s with a coefficient of variation of 55.4 percent, indicating highly unstable performance. Additionally, glm-5.2 experienced severe transient drops, plunging to 46.55 token/s at 22:30 and 88.69 token/s at 00:00. The dataset contains 288 valid points across six models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 191.34 token/s, while nemotron-3-ultra was weakest at 33.9 token/s. The most operationally significant volatility is deepseek-v4-flash, which swung sharply between 21.36 and 142.82 token/s with a 59.9 coefficient of variation, indicating highly unstable performance. Additionally, glm-5.2 experienced a severe throughput drop to 46.55 token/s at 22:30 before recovering. Although dataset coverage is 100.0 percent with 288 valid points, the five-minute observation interval limits the ability to detect sub-five-minute microbursts or brief outages, meaning rapid transient degradations are not captured.

4-hour window · 288 points · 100.0% coverage

Over the four-hour window, glm-5.2 delivered the strongest average throughput at 194.24 token/s, while nemotron-3-ultra was weakest at 29.28 token/s. The most operationally significant volatility is deepseek-v4-flash, which swung sharply between 21.36 and 135.06 token/s with a coefficient of variation of 60.1 percent. Nemotron-3-ultra also showed extreme instability, dropping to 3.68 token/s before trending upward by 80.8 percent. The dataset records 100.0 percent coverage across all six models with 288 valid points, but the four-hour duration limits any missing-data assessment of longer-term capacity planning.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 196.78 token/s, while nemotron-3-ultra is the weakest at 23.55 token/s. The most operationally significant volatility comes from deepseek-v4-flash and nemotron-3-ultra, which exhibit extreme throughput swings; deepseek-v4-flash fluctuates between 12.02 and 135.06 token/s, and nemotron-3-ultra varies from 2.64 to 71.78 token/s. This level of instability creates highly unpredictable latency for affected workloads. The dataset shows 100.0 percent coverage with 288 valid points, meaning there are no missing-data limitations impacting this specific analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 193.53 token/s, while nemotron-3-ultra was weakest at 24.61 token/s. Operationally, deepseek-v4-flash exhibited severe volatility with a 50.8% coefficient of variation, swinging between 11.88 and 133.32 token/s. Additionally, nemotron-3-ultra experienced a sharp operational degradation, dropping from 44.21 token/s at 17:05 to a minimum of 2.64 token/s at 18:30, reflecting a -26.5% trend. The dataset contains 288 valid observations, achieving 100.0% coverage across all 48 expected samples per model, meaning there are no missing-data limitations to constrain this analysis.

4-hour window · 270 points · 93.8% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 186.2 token/s, while nemotron-3-ultra was weakest at 26.77 token/s. Operationally, nemotron-3-ultra exhibited severe volatility and a sharp degradation, dropping from 56.06 token/s at 16:20 to 2.64 token/s at 18:30. deepseek-v4-flash also showed instability, spiking to 159.21 token/s at 16:25 before crashing to 8.22 token/s at 16:50. This analysis is limited by missing data; each model recorded 45 valid samples out of an expected 48, resulting in 93.8 percent coverage.

4-hour window · 198 points · 68.8% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 181.14 token/s, while nemotron-3-ultra was the weakest at 31.55 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which experienced a severe downward trend of 25.8 percent, plummeting to a minimum of 2.64 token/s. Similarly, deepseek-v4-flash showed high instability with a 44.4 percent coefficient of variation and repeated sharp drops below 15 token/s. A key limitation is that the dataset contains only 198 valid points out of an expected 288, resulting in 68.8 percent coverage. This missing data prevents a complete assessment of the full period.

4-hour window · 126 points · 43.8% coverage

Over the four-hour window, glm-5.2 is the strongest model with an average throughput of 173.55 token/s, while nemotron-3-ultra is the weakest at 37.89 token/s. Operationally, deepseek-v4-flash exhibits the most significant volatility, dropping to extreme lows of 8.22 token/s and 11.88 token/s, yielding a high coefficient of variation of 41.4 percent. All models have exactly 21 samples each, resulting in a dataset coverage of only 43.8 percent. This missing-data limitation restricts visibility into the first 140 minutes of the period, meaning the calculated averages may not represent full operational capacity.

4-hour window · 90 points · 31.2% coverage

Across the rolling four-hour window, glm-5.2 is the strongest model by average throughput at 164.88 token/s, while nemotron-3-ultra is the weakest at 38.09 token/s. The most operationally significant volatility occurs in deepseek-v4-flash, which exhibits severe throughput instability; despite an average of 80.49 token/s, it drops to extreme lows of 8.22 token/s and 11.88 token/s. This evaluation is constrained by a major missing-data limitation. The dataset contains only 90 valid observations out of an expected 288, representing 31.2% coverage, leaving the majority of the period unmonitored.