Across the four-hour window, glm-5.2 is the strongest model by average throughput at 90.88 token/s, closely followed by deepseek-v4-flash at 91.45 token/s. The weakest is nemotron-3-ultra at 33.9 token/s. Operationally, gemma4:31b shows the most significant volatility, with a high coefficient of variation of 49.8 percent and throughput dropping from a maximum of 132.98 token/s to a minimum of 15.1 token/s. It also exhibits a steep downward trend of -31.9 percent over the period. The dataset has 100.0 percent coverage with 288 valid points, so there are no missing-data limitations affecting this analysis.
Hourly performance insights
Summaries of rolling four-hour performance data
Across the four-hour window, deepseek-v4-flash had the highest average throughput at 92.84 token/s, while nemotron-3-ultra was weakest at 38.09 token/s. Operationally, gemma4:31b showed the most severe volatility, with a 52.5% coefficient of variation, dropping from a peak of 166.2 token/s to a latest reading of 23.08 token/s. Deepseek-v4-pro also exhibited sharp instability, plunging to 12.2 token/s at 23:40. In contrast, glm-5.2 was the only model with a positive trend, rising by 5.7% to average 90.84 token/s. The dataset contains 288 valid points across all six models, yielding 100.0% coverage with no missing-data limitations.
Across the four-hour window, glm-5.2 had the strongest average output-token throughput at 98.5 token/s, while nemotron-3-ultra was weakest at 39.84 token/s. The most operationally significant volatility came from gemma4:31b, which dropped sharply from a peak of 169.69 token/s to a low of 27.07 token/s, yielding a high coefficient of variation of 47.2% and a downward trend of 41.7%. Deepseek-v4-flash was the only model showing positive trend growth at 4.8%. The dataset contains 288 valid observations across all six models, achieving 100.0% coverage with no missing-data limitations.
Across the four-hour window, gemma4:31b had the strongest average output-token throughput at 101.78 token/s, while nemotron-3-ultra was weakest at 41.72 token/s. The most operationally significant volatility was gemma4:31b’s sharp late-window decline, dropping from a peak of 169.69 token/s down to 27.07 token/s, yielding a negative trend of 22.1 percent. Conversely, deepseek-v4-flash improved by 18.3 percent, ending at 92.53 token/s. Dataset coverage is complete at 100.0 percent across all 288 valid observations, meaning there are no missing-data limitations affecting this specific interval.
Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 106.24 token/s, while nemotron-3-ultra is the weakest at 39.39 token/s. The most operationally significant volatility comes from gemma4:31b, which fluctuates heavily between 55.26 and 169.69 token/s, yielding a coefficient of variation of 32.7 percent. Additionally, deepseek-v4-flash experienced a severe single-interval drop to 19.63 token/s at 22:40 UTC. The dataset contains 288 valid observations, achieving 100.0 percent coverage across all six models with no missing-data limitations.
Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 103.45 token/s, while nemotron-3-ultra is the weakest at 38.92 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits a coefficient of variation of 38.1 percent and drops to a minimum of 10.01 token/s. Additionally, deepseek-v4-flash shows a notable downward trend, declining 19.1 percent over the period to a latest value of 110.94 token/s. The dataset contains 288 valid observations, achieving 100.0 percent coverage across all models, meaning there are no missing-data limitations affecting this analysis.
Across the four-hour window, deepseek-v4-flash achieved the strongest average throughput at 102.08 token/s, while nemotron-3-ultra was the weakest at 34.11 token/s. The most operationally significant trend is deepseek-v4-flash's 23.7% throughput decline, dropping from early highs above 150 token/s to 67.71 token/s at 21:00. Additionally, nemotron-3-ultra exhibited severe volatility, swinging between 10.01 and 69.21 token/s with a 43.1% coefficient of variation. The dataset contains 288 valid observations across six models, achieving 100.0% coverage with no missing-data limitations.
Across the four-hour window, deepseek-v4-flash is the strongest model by average throughput at 109.08 token/s, while nemotron-3-ultra is the weakest at 29.36 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which has a coefficient of variation of 50.0% and throughput swings from 5.25 to 64.45 token/s. Deepseek-v4-flash also shows notable volatility, ranging from 41.53 to 194.47 token/s. The dataset includes 288 valid points across six models, achieving 100.0% coverage with no missing-data limitation over the expected 48 samples per model.
Across the four-hour window, deepseek-v4-flash is the strongest model with an average throughput of 115.67 token/s, while nemotron-3-ultra is the weakest at 27.13 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits a coefficient of variation of 53.1 percent and drops to a minimum of 5.25 token/s. Conversely, minimax-m3 shows a notable upward trend, increasing by 31.6 percent over the period to reach a late peak of 89.36 token/s. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, deepseek-v4-flash is the strongest model with an average throughput of 114.64 token/s, while nemotron-3-ultra is the weakest at 26.33 token/s. The most operationally significant volatility appears in minimax-m3, which drops to 6.55 token/s at 16:00 UTC before recovering. Nemotron-3-ultra also shows extreme instability, fluctuating between 5.25 and 60.53 token/s with a 24.3 percent downward trend. The dataset has 100.0 percent coverage with 288 valid points across six models, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, deepseek-v4-flash is the strongest model with an average throughput of 110.76 token/s, while nemotron-3-ultra is the weakest at 27.65 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits a coefficient of variation of 54.9 percent and a negative trend of 30.3 percent, dropping to a minimum of 5.25 token/s. Similarly, minimax-m3 shows severe instability with sudden drops to 6.55 token/s and a negative trend of 22.6 percent. The dataset includes 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, deepseek-v4-flash had the strongest average throughput at 107.2 token/s, while nemotron-3-ultra was weakest at 33.27 token/s. The most operationally significant volatility came from deepseek-v4-flash, which dropped to a minimum of 9.23 token/s at 12:35 despite reaching a maximum of 165.72 token/s at 12:50. glm-5.2 showed the largest upward trend, increasing by 20.1 percent to a latest value of 118.23 token/s. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, deepseek-v4-flash had the strongest average output-token throughput at 101.4 token/s, while nemotron-3-ultra was weakest at 32.38 token/s. The most operationally significant volatility came from deepseek-v4-flash, which ranged from a low of 9.23 token/s to a high of 165.72 token/s, reflecting high instability despite its leading average. Nemotron-3-ultra also showed notable volatility, dropping to 5.86 token/s. The dataset contains 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitations. All models maintained full sample counts, providing complete visibility into throughput performance for this period.
Across the four-hour window, deepseek-v4-flash had the highest average throughput at 102.2 token/s, while nemotron-3-ultra was the weakest at 33.31 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which dropped sharply from 143.88 token/s at 09:50 to 13.6 token/s at 10:00, and later hit a minimum of 9.23 token/s at 12:35. glm-5.2 also exhibited a severe drop to 11.16 token/s at 11:00. The dataset contains 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, deepseek-v4-flash had the highest average throughput at 103.02 token/s, while nemotron-3-ultra was the weakest at 33.36 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which dropped sharply from 143.88 token/s at 09:50 to 13.6 token/s at 10:00 before recovering to 112.54 token/s by 10:20. This model also exhibited the widest absolute throughput range, with a minimum of 13.6 token/s and a maximum of 147.15 token/s. The dataset contains 288 valid observations across all six models, achieving 100.0 percent coverage with no missing data limitations.
Across the four-hour window, deepseek-v4-flash achieved the highest average throughput at 104.88 token/s, while nemotron-3-ultra was the weakest at 33.02 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which dropped sharply from 143.88 token/s at 09:50 to 13.6 token/s at 10:00 before recovering to 112.54 token/s by 10:20. Similarly, glm-5.2 experienced a sudden decline to 15.24 token/s at 09:00 from 117.99 token/s at 08:55. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing data limitations.
Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 107.11 token/s, while nemotron-3-ultra is the weakest at 35.85 token/s. The most operationally significant volatility comes from deepseek-v4-flash, which maintains a high average of 106.18 token/s but drops sharply to a minimum of 13.6 token/s at the end of the period. Nemotron-3-ultra also shows notable instability with a high coefficient of variation of 44.8 percent and a downward trend of 18.3 percent. The dataset contains 288 valid observations with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, deepseek-v4-flash shows the strongest average throughput at 106.54 token/s, while nemotron-3-ultra is the weakest at 37.82 token/s. The most operationally significant volatility appears in gemma4:31b, which has a coefficient of variation of 36.7% and swings from a minimum of 29.56 token/s to a maximum of 167.19 token/s. Additionally, glm-5.2 exhibits a sharp late-window throughput drop, falling from 117.99 token/s at 08:55 to 15.24 token/s at 09:00. The dataset contains 288 valid points across six models, achieving 100.0% coverage with no missing-data limitations.
Across the four-hour window, gemma4:31b delivered the highest average throughput at 106.92 token/s, while nemotron-3-ultra was the weakest at 41.95 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which dropped sharply from 140.19 token/s at 07:50 to 15.7 token/s at 07:55 before recovering to 123.5 token/s at 08:00. glm-5.2 showed the strongest upward trend, increasing by 33.3 percent to a latest throughput of 116.56 token/s. The dataset includes 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, deepseek-v4-flash achieved the strongest average throughput at 103.84 token/s, while nemotron-3-ultra was the weakest at 45.14 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from a peak of 216.98 token/s down to 13.59 token/s at 04:30, reflecting a coefficient of variation of 53.4 percent. This dataset contains 288 valid observations across six models, corresponding to exactly 100.0 percent coverage of the 48 expected samples per model, meaning there are no missing-data limitations to report.
Across the four-hour window, glm-5.2 had the highest average output-token throughput at 122.38 token/s, while nemotron-3-ultra was weakest at 46.77 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from a peak of 223.93 token/s at 02:25 to a low of 13.59 token/s at 04:30, reflecting a 60.7 percent downward trend. In contrast, deepseek-v4-flash maintained the most stable performance, averaging 106.89 token/s with a coefficient of variation of 20.7 percent. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, glm-5.2 had the strongest average throughput at 153.49 token/s, while nemotron-3-ultra was weakest at 47.92 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from over 200 token/s early in the window to a minimum of 13.59 token/s at 04:30, reflecting a 43.8 percent downward trend. In contrast, minimax-m3 remained stable with a coefficient of variation of 18.1 percent and an average of 63.38 token/s. The dataset contains 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitation.
Across the 288 valid observations, glm-5.2 is the strongest model with an average throughput of 186.51 token/s, while nemotron-3-ultra is the weakest at 46.66 token/s. The most operationally significant trend is the sharp late-window throughput degradation in glm-5.2, which dropped from 216.98 token/s at 03:15 to a minimum of 58.75 token/s by 04:00, yielding a -10.9% period trend. Deepseek-v4-pro also exhibited high volatility, spiking between 123.03 token/s and 26.57 token/s. The dataset has 100.0% coverage with no missing-data limitation across the expected 48 samples per model.
Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 196.91 token/s, while nemotron-3-ultra is the weakest at 41.71 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits a coefficient of variation of 42.4% and drops to a minimum of 5.25 token/s, despite a 34.5% upward trend. Conversely, glm-5.2 maintains the highest stability with a coefficient of variation of just 10.2%. The dataset includes 288 valid observations, achieving 100.0% coverage across all 48 expected samples per model, meaning there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 195.03 token/s, while nemotron-3-ultra was weakest at 39.86 token/s. The most operationally significant volatility came from nemotron-3-ultra, which dropped to 5.25 token/s at 23:30 and exhibited a coefficient of variation of 43.9 percent. Deepseek-v4-pro and deepseek-v4-flash also experienced severe transient drops to 8.78 and 14.69 token/s, respectively. Dataset coverage is 100.0 percent with 288 valid points, meaning no missing-data limitation affects this specific analysis window.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 188.70 token/s, while nemotron-3-ultra was weakest at 32.04 token/s. Operationally, nemotron-3-ultra exhibited extreme volatility with a 59.3% coefficient of variation, swinging between 4.42 and 67.56 token/s. Additionally, deepseek-v4-pro and deepseek-v4-flash experienced severe transient drops, plummeting to 8.78 and 14.69 token/s respectively. The dataset includes 288 valid observations across six models, achieving 100.0% coverage with no missing data limitations.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 186.23 token/s, while nemotron-3-ultra was the weakest at 32.32 token/s. Operationally, nemotron-3-ultra exhibited extreme volatility with a 61.5% coefficient of variation, frequently dropping below 10 token/s. Additionally, deepseek-v4-pro and deepseek-v4-flash experienced severe transient slowdowns, plunging to 8.78 token/s and 14.69 token/s respectively. Although the dataset reports 100.0% coverage across all 288 valid samples, the five-minute granularity limits the ability to detect sub-interval micro-outages or pinpoint the exact duration of these sharp throughput drops.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 178.42 token/s, while nemotron-3-ultra was the weakest at 33.78 token/s. Operationally, nemotron-3-ultra showed severe volatility with a 58.4% coefficient of variation and a sharp downward trend, dropping from 45.26 token/s at 19:05 to 10.34 token/s by 23:00. Deepseek-v4-flash also exhibited instability, plunging to 14.69 token/s at 22:25. Although dataset coverage is 100.0% with 288 valid points, the four-hour window limits visibility into diurnal load patterns.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 175.08 token/s, while nemotron-3-ultra was weakest at 33.08 token/s. Operationally, nemotron-3-ultra showed extreme volatility, swinging between 4.29 and 68.67 token/s with a coefficient of variation of 57.6%, and its throughput trended downward by 21.2%. In contrast, minimax-m3 maintained the most stable performance, averaging 66.99 token/s with a coefficient of variation of just 17.1%. The dataset includes 288 valid observations, achieving 100.0% coverage across all six models with no missing data limitations.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 172.69 token/s, while nemotron-3-ultra was the weakest at 37.07 token/s. Operationally, deepseek-v4-pro shows the most significant degrading trend, dropping 16.1 percent to a latest reading of 64.85 token/s. Meanwhile, gemma4:31b exhibited extreme volatility, swinging between a minimum of 32.12 token/s and a maximum of 178.8 token/s. The dataset contains 288 valid observations across six models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this operational review.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 168.37 token/s, while nemotron-3-ultra was weakest at 31.61 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which logged a 57.1% coefficient of variation and a severe minimum of 2.26 token/s, indicating highly unstable generation. Conversely, deepseek-v4-pro exhibited a notable downward trend, dropping 10.4% to an average of 78.81 token/s. The dataset contains 288 valid points across all six models with 100.0% coverage, meaning there are no missing-data limitations affecting this analysis.
Over the four-hour window, glm-5.2 delivered the strongest average throughput at 172.86 token/s, while nemotron-3-ultra was weakest at 25.1 token/s. Operationally, nemotron-3-ultra showed extreme volatility, crashing to 1.67 token/s around 15:35 before recovering to 72.03 token/s at 16:45. Deepseek-v4-pro also exhibited sharp early swings, dropping to 11.73 token/s at 15:05. All six models achieved 100.0 percent coverage across 288 valid observations, meaning no missing-data limitation affects this specific dataset. However, the dataset lacks per-request concurrency details, limiting deeper capacity analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 177.87 token/s, while nemotron-3-ultra was weakest at 21.25 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which exhibited an 84.7% coefficient of variation and dropped to a minimum of 1.67 token/s. Deepseek-v4-pro also showed notable instability, falling to 11.73 token/s. The dataset contains 288 valid observations across six models, achieving 100.0% coverage. Because the dataset spans only four hours, it lacks longer-duration context, limiting the ability to assess diurnal patterns or sustained capacity degradation.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 183.76 token/s, while nemotron-3-ultra was weakest at 20.1 token/s. Operationally, nemotron-3-ultra exhibited the most significant volatility and degradation, plunging from a 72.07 token/s peak to sustained sub-3 token/s lows with a -39.8 percent trend. Although the dataset reports 100.0 percent coverage across all 288 valid point counts, the five-minute granularity limits the ability to detect sub-interval latency spikes or brief throttling events that would impact real-time user experience.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 184.47 token/s, while nemotron-3-ultra was weakest at 23.35 token/s. The most operationally significant volatility is nemotron-3-ultra's severe degradation, dropping from a 72.07 token/s peak to a 1.67 token/s minimum with a 85.4% coefficient of variation and a -46.4% trend. Deepseek-v4-flash also showed instability, plunging to 7.06 token/s at 13:00. Dataset coverage is 100.0% across all 288 valid points, meaning there are no missing-data limitations impacting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 186.74 token/s, while nemotron-3-ultra was the weakest at 25.67 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which exhibited extreme instability with a coefficient of variation of 78.0% and throughput dropping to a minimum of 2.3 token/s. This level of volatility indicates highly erratic performance degradation compared to the more stable models. The dataset is complete with 100.0% coverage and no missing data limitations across all 48 expected samples per model.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 181.76 token/s, while nemotron-3-ultra was weakest at 27.42 token/s. The most operationally significant volatility occurred in deepseek-v4-pro, which trended down 18.1 percent and dropped to a minimum of 13.99 token/s. Deepseek-v4-flash also showed severe instability, plummeting to 7.06 token/s at 13:00. Dataset coverage is complete with 288 valid observations, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 182.1 token/s, while nemotron-3-ultra was weakest at 29.76 token/s. The most operationally significant volatility is nemotron-3-ultra's severe late-morning throughput collapse, dropping from 50.35 token/s at 11:15 to 3.23 token/s at 12:05, driving its 60.5% coefficient of variation. Additionally, deepseek-v4-flash and deepseek-v4-pro experienced sharp final-interval drops to 7.06 token/s and 13.99 token/s, respectively. The dataset is complete with 100.0% coverage and 288 valid points, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 185.4 token/s, while nemotron-3-ultra was weakest at 31.63 token/s. The most operationally significant volatility is nemotron-3-ultra's 64.9% coefficient of variation, including a drop to 0.4 token/s at 08:35 and a declining trend of -37.1% to 4.92 token/s by 12:00. Although dataset coverage is 100.0% with 288 valid points, the five-minute sampling interval limits the detection of sub-five-minute micro-outages or transient latency spikes.
Across the four-hour dataset, glm-5.2 delivered the strongest average throughput at 177.71 token/s, while nemotron-3-ultra was weakest at 35.06 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which dropped to 0.4 token/s at 08:35, reflecting severe latency instability. deepseek-v4-flash showed the largest negative trend at -10.6 percent. Although coverage is 100 percent across all models, the dataset is limited by the absence of request concurrency metadata, preventing isolation of tenant load effects on these throughput drops.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 182.69 token/s, while nemotron-3-ultra was weakest at 37.46 token/s. Operationally, nemotron-3-ultra exhibited the most severe volatility, plunging to 0.4 token/s at 08:35, which indicates intermittent service disruptions. Conversely, minimax-m3 showed a notable declining trend, dropping 11.1 percent to a latest throughput of 25.93 token/s. Although the dataset reports 100.0 percent coverage across all 288 valid point counts, the five-minute observation interval lacks the granularity needed to determine if these extreme throughput drops represent total outages or momentary throttling.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 181.57 token/s, while nemotron-3-ultra was weakest at 35.11 token/s. The most operationally significant volatility is nemotron-3-ultra's severe throughput collapse between 08:30 and 08:35 UTC, where output plummeted from 1.47 token/s to 0.4 token/s before recovering to 28.18 token/s. This dataset contains no missing-data limitation, as all six models achieved 100.0% coverage across their expected 48 samples.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 180.86 token/s, while nemotron-3-ultra was weakest at 30.78 token/s. The most operationally significant volatility appears in glm-5.2, which despite its high average suffered sharp five-minute drops to 39.01 token/s and 58.29 token/s. Nemotron-3-ultra also showed extreme instability, ranging from 4.16 to 69.35 token/s with a 49.2% coefficient of variation. All six models achieved 100% coverage with 288 valid points, so there is no missing-data limitation in this dataset.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 187.27 token/s, while nemotron-3-ultra was weakest at 33.11 token/s. Operationally, deepseek-v4-flash and nemotron-3-ultra showed severe volatility, with coefficients of variation reaching 35.9% and 52.5% respectively, driven by sharp throughput drops to 22.35 token/s and 4.16 token/s. Although dataset coverage is 100.0%, the analysis is limited by the absence of concurrent request counts, preventing differentiation between model-side latency spikes and user-driven load fluctuations.
Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 185.69 token/s, while nemotron-3-ultra was the weakest at 34.49 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which dropped sharply from a 69.73 token/s peak to a 4.16 token/s minimum, reflecting a 41.5 percent downward trend. deepseek-v4-pro also showed notable instability, falling to 20.18 token/s. The dataset contains 288 valid observations across six models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 190.22 token/s, while nemotron-3-ultra was weakest at 37.63 token/s. The most operationally significant volatility appears in deepseek-v4-flash, which swung sharply between 18.33 and 152.68 token/s with a coefficient of variation of 46.4 percent, indicating highly unstable response times. Additionally, deepseek-v4-pro exhibited a notable downward trend, dropping 19.0 percent over the period and hitting a low of 20.18 token/s. Although coverage is 100.0 percent, the dataset lacks concurrent request volume data, limiting the ability to determine if throughput drops stem from capacity saturation or internal model inefficiencies.
Over the four-hour window, glm-5.2 had the strongest average throughput at 187.89 token/s, while nemotron-3-ultra was weakest at 41.06 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which exhibited extreme swings between 155.36 token/s and 18.33 token/s, resulting in a coefficient of variation of 52.5%. This erratic performance complicates capacity planning. Although dataset coverage is 100.0% with no missing samples, the analysis is limited by the brief four-hour duration, which prevents assessing diurnal patterns or long-term degradation.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 187.87 token/s, while nemotron-3-ultra was weakest at 38.29 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which exhibited extreme swings between 155.36 token/s and 18.33 token/s, yielding a coefficient of variation of 50.6 percent. This erratic performance creates unpredictable latency for end users. Although dataset coverage is marked as 100.0 percent with 288 valid points, the five-minute sampling interval lacks sub-minute resolution, limiting the ability to detect brief micro-spike degradations.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 186.96 token/s, while nemotron-3-ultra was weakest at 38.17 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which exhibited extreme swings between 19.96 and 155.36 token/s, reflecting a 51.2% coefficient of variation. Similarly, nemotron-3-ultra suffered severe micro-drops, hitting a low of 1.9 token/s. Although the dataset reports 100.0% coverage across all models, the five-minute sampling interval lacks the granularity needed to isolate the root causes of these sharp transient drops.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 185.62 token/s, while nemotron-3-ultra was weakest at 35.94 token/s. The most operationally significant volatility is deepseek-v4-flash, which exhibited extreme swings between 19.96 and 155.36 token/s with a coefficient of variation of 55.4 percent, indicating highly unstable performance. Additionally, glm-5.2 experienced severe transient drops, plunging to 46.55 token/s at 22:30 and 88.69 token/s at 00:00. The dataset contains 288 valid points across six models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this analysis.