Across the four-hour window, deepseek-v4-flash achieved the strongest average throughput at 103.84 token/s, while nemotron-3-ultra was the weakest at 45.14 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from a peak of 216.98 token/s down to 13.59 token/s at 04:30, reflecting a coefficient of variation of 53.4 percent. This dataset contains 288 valid observations across six models, corresponding to exactly 100.0 percent coverage of the 48 expected samples per model, meaning there are no missing-data limitations to report.
Hourly performance insights
GLM-5.2 summaries of rolling four-hour performance data
Across the four-hour window, glm-5.2 had the highest average output-token throughput at 122.38 token/s, while nemotron-3-ultra was weakest at 46.77 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from a peak of 223.93 token/s at 02:25 to a low of 13.59 token/s at 04:30, reflecting a 60.7 percent downward trend. In contrast, deepseek-v4-flash maintained the most stable performance, averaging 106.89 token/s with a coefficient of variation of 20.7 percent. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, glm-5.2 had the strongest average throughput at 153.49 token/s, while nemotron-3-ultra was weakest at 47.92 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from over 200 token/s early in the window to a minimum of 13.59 token/s at 04:30, reflecting a 43.8 percent downward trend. In contrast, minimax-m3 remained stable with a coefficient of variation of 18.1 percent and an average of 63.38 token/s. The dataset contains 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitation.
Across the 288 valid observations, glm-5.2 is the strongest model with an average throughput of 186.51 token/s, while nemotron-3-ultra is the weakest at 46.66 token/s. The most operationally significant trend is the sharp late-window throughput degradation in glm-5.2, which dropped from 216.98 token/s at 03:15 to a minimum of 58.75 token/s by 04:00, yielding a -10.9% period trend. Deepseek-v4-pro also exhibited high volatility, spiking between 123.03 token/s and 26.57 token/s. The dataset has 100.0% coverage with no missing-data limitation across the expected 48 samples per model.
Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 196.91 token/s, while nemotron-3-ultra is the weakest at 41.71 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits a coefficient of variation of 42.4% and drops to a minimum of 5.25 token/s, despite a 34.5% upward trend. Conversely, glm-5.2 maintains the highest stability with a coefficient of variation of just 10.2%. The dataset includes 288 valid observations, achieving 100.0% coverage across all 48 expected samples per model, meaning there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 195.03 token/s, while nemotron-3-ultra was weakest at 39.86 token/s. The most operationally significant volatility came from nemotron-3-ultra, which dropped to 5.25 token/s at 23:30 and exhibited a coefficient of variation of 43.9 percent. Deepseek-v4-pro and deepseek-v4-flash also experienced severe transient drops to 8.78 and 14.69 token/s, respectively. Dataset coverage is 100.0 percent with 288 valid points, meaning no missing-data limitation affects this specific analysis window.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 188.70 token/s, while nemotron-3-ultra was weakest at 32.04 token/s. Operationally, nemotron-3-ultra exhibited extreme volatility with a 59.3% coefficient of variation, swinging between 4.42 and 67.56 token/s. Additionally, deepseek-v4-pro and deepseek-v4-flash experienced severe transient drops, plummeting to 8.78 and 14.69 token/s respectively. The dataset includes 288 valid observations across six models, achieving 100.0% coverage with no missing data limitations.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 186.23 token/s, while nemotron-3-ultra was the weakest at 32.32 token/s. Operationally, nemotron-3-ultra exhibited extreme volatility with a 61.5% coefficient of variation, frequently dropping below 10 token/s. Additionally, deepseek-v4-pro and deepseek-v4-flash experienced severe transient slowdowns, plunging to 8.78 token/s and 14.69 token/s respectively. Although the dataset reports 100.0% coverage across all 288 valid samples, the five-minute granularity limits the ability to detect sub-interval micro-outages or pinpoint the exact duration of these sharp throughput drops.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 178.42 token/s, while nemotron-3-ultra was the weakest at 33.78 token/s. Operationally, nemotron-3-ultra showed severe volatility with a 58.4% coefficient of variation and a sharp downward trend, dropping from 45.26 token/s at 19:05 to 10.34 token/s by 23:00. Deepseek-v4-flash also exhibited instability, plunging to 14.69 token/s at 22:25. Although dataset coverage is 100.0% with 288 valid points, the four-hour window limits visibility into diurnal load patterns.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 175.08 token/s, while nemotron-3-ultra was weakest at 33.08 token/s. Operationally, nemotron-3-ultra showed extreme volatility, swinging between 4.29 and 68.67 token/s with a coefficient of variation of 57.6%, and its throughput trended downward by 21.2%. In contrast, minimax-m3 maintained the most stable performance, averaging 66.99 token/s with a coefficient of variation of just 17.1%. The dataset includes 288 valid observations, achieving 100.0% coverage across all six models with no missing data limitations.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 172.69 token/s, while nemotron-3-ultra was the weakest at 37.07 token/s. Operationally, deepseek-v4-pro shows the most significant degrading trend, dropping 16.1 percent to a latest reading of 64.85 token/s. Meanwhile, gemma4:31b exhibited extreme volatility, swinging between a minimum of 32.12 token/s and a maximum of 178.8 token/s. The dataset contains 288 valid observations across six models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this operational review.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 168.37 token/s, while nemotron-3-ultra was weakest at 31.61 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which logged a 57.1% coefficient of variation and a severe minimum of 2.26 token/s, indicating highly unstable generation. Conversely, deepseek-v4-pro exhibited a notable downward trend, dropping 10.4% to an average of 78.81 token/s. The dataset contains 288 valid points across all six models with 100.0% coverage, meaning there are no missing-data limitations affecting this analysis.
Over the four-hour window, glm-5.2 delivered the strongest average throughput at 172.86 token/s, while nemotron-3-ultra was weakest at 25.1 token/s. Operationally, nemotron-3-ultra showed extreme volatility, crashing to 1.67 token/s around 15:35 before recovering to 72.03 token/s at 16:45. Deepseek-v4-pro also exhibited sharp early swings, dropping to 11.73 token/s at 15:05. All six models achieved 100.0 percent coverage across 288 valid observations, meaning no missing-data limitation affects this specific dataset. However, the dataset lacks per-request concurrency details, limiting deeper capacity analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 177.87 token/s, while nemotron-3-ultra was weakest at 21.25 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which exhibited an 84.7% coefficient of variation and dropped to a minimum of 1.67 token/s. Deepseek-v4-pro also showed notable instability, falling to 11.73 token/s. The dataset contains 288 valid observations across six models, achieving 100.0% coverage. Because the dataset spans only four hours, it lacks longer-duration context, limiting the ability to assess diurnal patterns or sustained capacity degradation.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 183.76 token/s, while nemotron-3-ultra was weakest at 20.1 token/s. Operationally, nemotron-3-ultra exhibited the most significant volatility and degradation, plunging from a 72.07 token/s peak to sustained sub-3 token/s lows with a -39.8 percent trend. Although the dataset reports 100.0 percent coverage across all 288 valid point counts, the five-minute granularity limits the ability to detect sub-interval latency spikes or brief throttling events that would impact real-time user experience.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 184.47 token/s, while nemotron-3-ultra was weakest at 23.35 token/s. The most operationally significant volatility is nemotron-3-ultra's severe degradation, dropping from a 72.07 token/s peak to a 1.67 token/s minimum with a 85.4% coefficient of variation and a -46.4% trend. Deepseek-v4-flash also showed instability, plunging to 7.06 token/s at 13:00. Dataset coverage is 100.0% across all 288 valid points, meaning there are no missing-data limitations impacting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 186.74 token/s, while nemotron-3-ultra was the weakest at 25.67 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which exhibited extreme instability with a coefficient of variation of 78.0% and throughput dropping to a minimum of 2.3 token/s. This level of volatility indicates highly erratic performance degradation compared to the more stable models. The dataset is complete with 100.0% coverage and no missing data limitations across all 48 expected samples per model.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 181.76 token/s, while nemotron-3-ultra was weakest at 27.42 token/s. The most operationally significant volatility occurred in deepseek-v4-pro, which trended down 18.1 percent and dropped to a minimum of 13.99 token/s. Deepseek-v4-flash also showed severe instability, plummeting to 7.06 token/s at 13:00. Dataset coverage is complete with 288 valid observations, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 182.1 token/s, while nemotron-3-ultra was weakest at 29.76 token/s. The most operationally significant volatility is nemotron-3-ultra's severe late-morning throughput collapse, dropping from 50.35 token/s at 11:15 to 3.23 token/s at 12:05, driving its 60.5% coefficient of variation. Additionally, deepseek-v4-flash and deepseek-v4-pro experienced sharp final-interval drops to 7.06 token/s and 13.99 token/s, respectively. The dataset is complete with 100.0% coverage and 288 valid points, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 185.4 token/s, while nemotron-3-ultra was weakest at 31.63 token/s. The most operationally significant volatility is nemotron-3-ultra's 64.9% coefficient of variation, including a drop to 0.4 token/s at 08:35 and a declining trend of -37.1% to 4.92 token/s by 12:00. Although dataset coverage is 100.0% with 288 valid points, the five-minute sampling interval limits the detection of sub-five-minute micro-outages or transient latency spikes.
Across the four-hour dataset, glm-5.2 delivered the strongest average throughput at 177.71 token/s, while nemotron-3-ultra was weakest at 35.06 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which dropped to 0.4 token/s at 08:35, reflecting severe latency instability. deepseek-v4-flash showed the largest negative trend at -10.6 percent. Although coverage is 100 percent across all models, the dataset is limited by the absence of request concurrency metadata, preventing isolation of tenant load effects on these throughput drops.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 182.69 token/s, while nemotron-3-ultra was weakest at 37.46 token/s. Operationally, nemotron-3-ultra exhibited the most severe volatility, plunging to 0.4 token/s at 08:35, which indicates intermittent service disruptions. Conversely, minimax-m3 showed a notable declining trend, dropping 11.1 percent to a latest throughput of 25.93 token/s. Although the dataset reports 100.0 percent coverage across all 288 valid point counts, the five-minute observation interval lacks the granularity needed to determine if these extreme throughput drops represent total outages or momentary throttling.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 181.57 token/s, while nemotron-3-ultra was weakest at 35.11 token/s. The most operationally significant volatility is nemotron-3-ultra's severe throughput collapse between 08:30 and 08:35 UTC, where output plummeted from 1.47 token/s to 0.4 token/s before recovering to 28.18 token/s. This dataset contains no missing-data limitation, as all six models achieved 100.0% coverage across their expected 48 samples.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 180.86 token/s, while nemotron-3-ultra was weakest at 30.78 token/s. The most operationally significant volatility appears in glm-5.2, which despite its high average suffered sharp five-minute drops to 39.01 token/s and 58.29 token/s. Nemotron-3-ultra also showed extreme instability, ranging from 4.16 to 69.35 token/s with a 49.2% coefficient of variation. All six models achieved 100% coverage with 288 valid points, so there is no missing-data limitation in this dataset.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 187.27 token/s, while nemotron-3-ultra was weakest at 33.11 token/s. Operationally, deepseek-v4-flash and nemotron-3-ultra showed severe volatility, with coefficients of variation reaching 35.9% and 52.5% respectively, driven by sharp throughput drops to 22.35 token/s and 4.16 token/s. Although dataset coverage is 100.0%, the analysis is limited by the absence of concurrent request counts, preventing differentiation between model-side latency spikes and user-driven load fluctuations.
Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 185.69 token/s, while nemotron-3-ultra was the weakest at 34.49 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which dropped sharply from a 69.73 token/s peak to a 4.16 token/s minimum, reflecting a 41.5 percent downward trend. deepseek-v4-pro also showed notable instability, falling to 20.18 token/s. The dataset contains 288 valid observations across six models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 190.22 token/s, while nemotron-3-ultra was weakest at 37.63 token/s. The most operationally significant volatility appears in deepseek-v4-flash, which swung sharply between 18.33 and 152.68 token/s with a coefficient of variation of 46.4 percent, indicating highly unstable response times. Additionally, deepseek-v4-pro exhibited a notable downward trend, dropping 19.0 percent over the period and hitting a low of 20.18 token/s. Although coverage is 100.0 percent, the dataset lacks concurrent request volume data, limiting the ability to determine if throughput drops stem from capacity saturation or internal model inefficiencies.
Over the four-hour window, glm-5.2 had the strongest average throughput at 187.89 token/s, while nemotron-3-ultra was weakest at 41.06 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which exhibited extreme swings between 155.36 token/s and 18.33 token/s, resulting in a coefficient of variation of 52.5%. This erratic performance complicates capacity planning. Although dataset coverage is 100.0% with no missing samples, the analysis is limited by the brief four-hour duration, which prevents assessing diurnal patterns or long-term degradation.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 187.87 token/s, while nemotron-3-ultra was weakest at 38.29 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which exhibited extreme swings between 155.36 token/s and 18.33 token/s, yielding a coefficient of variation of 50.6 percent. This erratic performance creates unpredictable latency for end users. Although dataset coverage is marked as 100.0 percent with 288 valid points, the five-minute sampling interval lacks sub-minute resolution, limiting the ability to detect brief micro-spike degradations.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 186.96 token/s, while nemotron-3-ultra was weakest at 38.17 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which exhibited extreme swings between 19.96 and 155.36 token/s, reflecting a 51.2% coefficient of variation. Similarly, nemotron-3-ultra suffered severe micro-drops, hitting a low of 1.9 token/s. Although the dataset reports 100.0% coverage across all models, the five-minute sampling interval lacks the granularity needed to isolate the root causes of these sharp transient drops.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 185.62 token/s, while nemotron-3-ultra was weakest at 35.94 token/s. The most operationally significant volatility is deepseek-v4-flash, which exhibited extreme swings between 19.96 and 155.36 token/s with a coefficient of variation of 55.4 percent, indicating highly unstable performance. Additionally, glm-5.2 experienced severe transient drops, plunging to 46.55 token/s at 22:30 and 88.69 token/s at 00:00. The dataset contains 288 valid points across six models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 191.34 token/s, while nemotron-3-ultra was weakest at 33.9 token/s. The most operationally significant volatility is deepseek-v4-flash, which swung sharply between 21.36 and 142.82 token/s with a 59.9 coefficient of variation, indicating highly unstable performance. Additionally, glm-5.2 experienced a severe throughput drop to 46.55 token/s at 22:30 before recovering. Although dataset coverage is 100.0 percent with 288 valid points, the five-minute observation interval limits the ability to detect sub-five-minute microbursts or brief outages, meaning rapid transient degradations are not captured.
Over the four-hour window, glm-5.2 delivered the strongest average throughput at 194.24 token/s, while nemotron-3-ultra was weakest at 29.28 token/s. The most operationally significant volatility is deepseek-v4-flash, which swung sharply between 21.36 and 135.06 token/s with a coefficient of variation of 60.1 percent. Nemotron-3-ultra also showed extreme instability, dropping to 3.68 token/s before trending upward by 80.8 percent. The dataset records 100.0 percent coverage across all six models with 288 valid points, but the four-hour duration limits any missing-data assessment of longer-term capacity planning.
Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 196.78 token/s, while nemotron-3-ultra is the weakest at 23.55 token/s. The most operationally significant volatility comes from deepseek-v4-flash and nemotron-3-ultra, which exhibit extreme throughput swings; deepseek-v4-flash fluctuates between 12.02 and 135.06 token/s, and nemotron-3-ultra varies from 2.64 to 71.78 token/s. This level of instability creates highly unpredictable latency for affected workloads. The dataset shows 100.0 percent coverage with 288 valid points, meaning there are no missing-data limitations impacting this specific analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 193.53 token/s, while nemotron-3-ultra was weakest at 24.61 token/s. Operationally, deepseek-v4-flash exhibited severe volatility with a 50.8% coefficient of variation, swinging between 11.88 and 133.32 token/s. Additionally, nemotron-3-ultra experienced a sharp operational degradation, dropping from 44.21 token/s at 17:05 to a minimum of 2.64 token/s at 18:30, reflecting a -26.5% trend. The dataset contains 288 valid observations, achieving 100.0% coverage across all 48 expected samples per model, meaning there are no missing-data limitations to constrain this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 186.2 token/s, while nemotron-3-ultra was weakest at 26.77 token/s. Operationally, nemotron-3-ultra exhibited severe volatility and a sharp degradation, dropping from 56.06 token/s at 16:20 to 2.64 token/s at 18:30. deepseek-v4-flash also showed instability, spiking to 159.21 token/s at 16:25 before crashing to 8.22 token/s at 16:50. This analysis is limited by missing data; each model recorded 45 valid samples out of an expected 48, resulting in 93.8 percent coverage.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 181.14 token/s, while nemotron-3-ultra was the weakest at 31.55 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which experienced a severe downward trend of 25.8 percent, plummeting to a minimum of 2.64 token/s. Similarly, deepseek-v4-flash showed high instability with a 44.4 percent coefficient of variation and repeated sharp drops below 15 token/s. A key limitation is that the dataset contains only 198 valid points out of an expected 288, resulting in 68.8 percent coverage. This missing data prevents a complete assessment of the full period.
Over the four-hour window, glm-5.2 is the strongest model with an average throughput of 173.55 token/s, while nemotron-3-ultra is the weakest at 37.89 token/s. Operationally, deepseek-v4-flash exhibits the most significant volatility, dropping to extreme lows of 8.22 token/s and 11.88 token/s, yielding a high coefficient of variation of 41.4 percent. All models have exactly 21 samples each, resulting in a dataset coverage of only 43.8 percent. This missing-data limitation restricts visibility into the first 140 minutes of the period, meaning the calculated averages may not represent full operational capacity.
Across the rolling four-hour window, glm-5.2 is the strongest model by average throughput at 164.88 token/s, while nemotron-3-ultra is the weakest at 38.09 token/s. The most operationally significant volatility occurs in deepseek-v4-flash, which exhibits severe throughput instability; despite an average of 80.49 token/s, it drops to extreme lows of 8.22 token/s and 11.88 token/s. This evaluation is constrained by a major missing-data limitation. The dataset contains only 90 valid observations out of an expected 288, representing 31.2% coverage, leaving the majority of the period unmonitored.