Across the four-hour window, deepseek-v4-flash had the strongest average output-token throughput at 86.16 token/s, while gemma4:31b was weakest at 23.88 token/s. The most operationally significant volatility came from nemotron-3-ultra, which dropped 45.8 percent over the period with a coefficient of variation of 76.2 percent, swinging between 104.14 token/s and 5.89 token/s. deepseek-v4-flash also showed high instability, plunging to 13.15 token/s before recovering. The dataset includes 288 valid observations across six models with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.
Hourly performance insights
Summaries of rolling four-hour performance data
Over the four-hour window, deepseek-v4-flash had the strongest average throughput at 88.92 token/s, while gemma4:31b was weakest at 30.81 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which had a coefficient of variation of 76.4 percent and throughput swings from 6.66 to 104.14 token/s. deepseek-v4-flash also showed sharp periodic drops, falling to 13.15 token/s at 09:00 despite a high p95 of 135.46 token/s. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitation.
Across the four-hour window, deepseek-v4-flash had the strongest average throughput at 101.75 token/s, while minimax-m3 was weakest at 36.09 token/s. Operationally, deepseek-v4-flash and nemotron-3-ultra showed severe volatility, with coefficients of variation of 39.1 percent and 73.8 percent respectively. Deepseek-v4-flash experienced steep drops to 17.64 token/s at 06:35 and 13.15 token/s at 09:00. Overall dataset coverage is 99.7 percent with 287 valid points, but gemma4:31b is missing one sample, yielding 47 points and 97.9 percent coverage, which slightly limits its trend accuracy.
Across the four-hour window, deepseek-v4-flash had the strongest average throughput at 111.4 token/s, while minimax-m3 was weakest at 39.45 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which swung from a low of 6.66 token/s to a high of 104.14 token/s with a coefficient of variation of 74.0 percent. deepseek-v4-flash also showed sharp instability, dropping to 17.64 token/s at 06:35 before recovering. This analysis is limited by a missing data gap for gemma4:31b, which recorded 47 samples instead of the expected 48, reducing its coverage to 97.9 percent compared to the 100 percent coverage of all other models.
Across the four-hour window, deepseek-v4-flash had the strongest average throughput at 111.72 token/s, while minimax-m3 was the weakest at 39.37 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which had a coefficient of variation of 72.0 percent and a downward trend of -23.9 percent, dropping from an early peak of 103.84 token/s to a low of 6.66 token/s. Overall dataset coverage reached 99.7 percent across 287 valid points, but a missing-data limitation affects gemma4:31b, which captured only 47 of 48 expected samples for 97.9 percent coverage.
Over the four-hour window, deepseek-v4-flash had the highest average throughput at 115.96 token/s, while minimax-m3 was the weakest at 38.25 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which swung from a low of 4.6 token/s to a high of 103.84 token/s with a coefficient of variation of 73.1 percent, indicating highly unstable output capacity. In contrast, glm-5.2 maintained steadier performance, averaging 92.02 token/s with a lower coefficient of variation of 26.6 percent. A missing-data limitation affects this review: gemma4:31b recorded only 47 samples instead of the expected 48, resulting in 97.9 percent coverage.
Across the four-hour window, deepseek-v4-flash is the strongest model with an average throughput of 113.36 token/s, while nemotron-3-ultra is the weakest at 36.36 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits a coefficient of variation of 67.9 percent and drops to a minimum of 4.6 token/s despite a maximum of 100.59 token/s. In contrast, deepseek-v4-flash maintains a stable trend with zero net change over the period. The dataset includes 288 valid observations, achieving 100.0 percent coverage across all six models, meaning there are no missing-data limitations affecting this analysis.
Across the four-hour window, deepseek-v4-flash had the highest average throughput at 121.68 token/s, while nemotron-3-ultra was the weakest at 35.74 token/s. The most operationally significant volatility came from nemotron-3-ultra, which exhibited extreme instability with a coefficient of variation of 69.3 percent and sharp drops to a minimum of 4.6 token/s. Additionally, gemma4:31b showed a notable downward trend, declining 24.8 percent over the period to a latest value of 24.24 token/s. The dataset includes 288 valid observations across six models, achieving complete coverage. However, a missing-data limitation exists because the provided points arrays omit the exact UTC dates for individual observations, relying solely on short timestamps.
Over the four-hour observation window, deepseek-v4-flash achieved the strongest average throughput at 122.11 token/s, while minimax-m3 was the weakest at 44.02 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which exhibited a 71.3% coefficient of variation, frequent steep drops to as low as 8.1 token/s, and a 42.8% negative trend, ending at 19.17 token/s. In contrast, deepseek-v4-pro maintained the most stable performance with a coefficient of variation of 18.4% and an average of 86.66 token/s. Dataset coverage is complete at 100.0% across all 288 valid samples, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, deepseek-v4-flash is the strongest model with an average throughput of 131.53 token/s, while minimax-m3 is the weakest at 46.94 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which has a coefficient of variation of 63.7% and a downward trend of -52.8%, dropping from a peak of 142.99 token/s to a low of 8.46 token/s. In contrast, deepseek-v4-flash maintains the most stable performance with a coefficient of variation of just 15.3%. The dataset includes 288 valid observations, achieving 100.0% coverage across all six models with no missing-data limitations.
Across the four-hour window, deepseek-v4-flash is the strongest model, averaging 133.41 token/s, while minimax-m3 is the weakest at 47.28 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which has a coefficient of variation of 55.9 percent and throughput drops as low as 8.46 token/s despite averaging 66.75 token/s. glm-5.2 shows a notable upward trend of 32.9 percent, peaking at 158.74 token/s. The dataset includes 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitation.
Across the four-hour window, deepseek-v4-flash achieved the strongest average throughput at 130.41 token/s, while minimax-m3 was the weakest at 46.21 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which exhibited extreme swings between 7.52 token/s and 142.99 token/s, yielding a coefficient of variation of 59.0 percent. In contrast, deepseek-v4-flash maintained the most stable performance with a coefficient of variation of just 17.0 percent. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage. Because the dataset contains no missing observations, there are no missing-data limitations affecting this specific four-hour analysis.
Across the four-hour observation window, deepseek-v4-flash is the strongest model with an average throughput of 128.19 token/s, while minimax-m3 is the weakest at 45.72 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which exhibits a coefficient of variation of 61.5 percent and throughput swings ranging from a minimum of 7.52 token/s to a maximum of 137.44 token/s. This extreme variability indicates highly unstable generation performance. The dataset includes complete coverage with 288 valid points across all models, resulting in no missing-data limitations for this specific interval.
Across the four-hour window, deepseek-v4-flash had the strongest average throughput at 118.76 token/s, while gemma4:31b was weakest at 36.46 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which ranged from 7.52 to 128.09 token/s with a coefficient of variation of 66.6%, indicating highly unstable generation speeds. Conversely, deepseek-v4-pro maintained the most consistent relative throughput, varying less overall despite a sharp final drop to 22.53 token/s. The dataset includes 288 valid observations across six models, achieving 100.0% coverage with no missing-data limitations. All expected samples were recorded, providing complete visibility into these throughput fluctuations.
Across the four-hour window, deepseek-v4-flash is the strongest model with an average throughput of 107.02 token/s, while gemma4:31b is the weakest at 27.17 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which has a coefficient of variation of 72.3 percent and throughput swings from 7.52 to 128.09 token/s. In contrast, deepseek-v4-pro offers the most stable performance, varying only 16.3 percent around its 87.36 token/s average. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage. Because the data is limited to this single rolling window, any missing-data limitation regarding broader diurnal patterns or prior day comparisons cannot be assessed.
Over the four-hour observation period, deepseek-v4-flash was the strongest model with an average throughput of 100.82 token/s, while gemma4:31b was the weakest at 21.5 token/s. The most operationally significant volatility appeared in gemma4:31b and nemotron-3-ultra, which exhibited high variability with coefficients of variation of 68.9% and 68.8%, respectively. This volatility included severe throughput drops, with gemma4:31b falling to a minimum of 6.01 token/s and nemotron-3-ultra dropping to 7.81 token/s. In contrast, deepseek-v4-pro maintained the most stable performance with a coefficient of variation of 22.9%. The dataset contains complete coverage with no missing-data limitations across all 288 valid observations.
Across the four-hour window, deepseek-v4-flash is the strongest model with an average throughput of 96.64 token/s, while gemma4:31b is the weakest at 17.39 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which has a coefficient of variation of 71.5% and throughput swings from 7.44 to 101.31 token/s. Similarly, deepseek-v4-flash experienced a severe drop to 17.69 token/s at 17:05 before recovering. The dataset includes 288 valid observations across six models, achieving 100.0% coverage. Because the dataset contains no missing data, there are no missing-data limitations affecting this specific throughput analysis.
Across the four-hour window, deepseek-v4-flash had the strongest average output-token throughput at 90.13 token/s, while gemma4:31b was the weakest at 13.18 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which dropped to a 4.51 token/s minimum and exhibited a 71.9% coefficient of variation, alongside a -12.5% downward trend. In contrast, deepseek-v4-flash showed a 10.8% upward trend, peaking at 117.41 token/s. The dataset contains 288 valid observations across six models, achieving 100.0% coverage with no missing-data limitations.
Across the four-hour window, deepseek-v4-flash is the strongest model with an average throughput of 89.7 token/s, while gemma4:31b is the weakest at 15.14 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which exhibits a coefficient of variation of 79.6 percent and oscillates between a minimum of 3.91 token/s and a maximum of 104.52 token/s. Deepseek-v4-flash also shows notable volatility, dropping to 15.79 token/s at 13:00 before stabilizing above 100 token/s later. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, deepseek-v4-flash had the strongest average output-token throughput at 85.13 token/s, while gemma4:31b was the weakest at 17.5 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which ranged from 3.91 to 104.52 token/s with a coefficient of variation of 85.9 percent, indicating highly unstable generation speeds. Conversely, deepseek-v4-pro showed the most stable relative performance with a coefficient of variation of 28.1 percent. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitation. All models reported the expected 48 samples, providing a complete view of current cloud throughput operations.
Across the four-hour window, deepseek-v4-flash had the highest average output-token throughput at 83.28 token/s, while gemma4:31b was the weakest at 21.12 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which swung between 3.91 and 104.52 token/s with a coefficient of variation of 83.9 percent, indicating highly unstable generation speeds. In contrast, glm-5.2 showed a strong upward trend, rising 27.6 percent to finish at 56.83 token/s. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, deepseek-v4-flash had the highest average output-token throughput at 89.33 token/s, while gemma4:31b was the weakest at 23.6 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which exhibited extreme swings between 3.91 and 114.46 token/s, yielding a coefficient of variation of 82.4 percent. Additionally, deepseek-v4-flash and deepseek-v4-pro both experienced sharp end-of-window drops, falling to 15.79 and 25.4 token/s respectively. The dataset includes 288 valid observations with 100.0 percent coverage, meaning there are no missing-data limitations affecting this analysis.
Over the four-hour window, deepseek-v4-flash had the strongest average throughput at 90.88 token/s, while gemma4:31b was weakest at 25.06 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which dropped 34.0 percent over the period to a low of 5.04 token/s, showing a coefficient of variation of 77.3 percent. In contrast, deepseek-v4-pro maintained the most stable performance with a coefficient of variation of 24.7 percent despite a brief drop to 20.06 token/s. The dataset includes 288 valid observations across six models with 100.0 percent coverage, so no missing-data limitation affects this specific interval.
Across the four-hour window, deepseek-v4-flash had the strongest average output-token throughput at 94.66 token/s, while gemma4:31b was the weakest at 27.25 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which had a coefficient of variation of 74.9 percent and a minimum of 4.27 token/s despite a maximum of 114.46 token/s. deepseek-v4-pro showed steadier performance with an average of 80.9 token/s and a lower coefficient of variation of 24.9 percent. Dataset coverage is complete at 100.0 percent across all 288 valid observations, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, deepseek-v4-flash is the strongest model with an average throughput of 97.44 token/s, while gemma4:31b is the weakest at 29.63 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits extreme swings between 4.27 and 104.5 token/s, resulting in a coefficient of variation of 68.9 percent. Additionally, glm-5.2 shows a sharp downward trend of 21.9 percent over the period, dropping from early highs above 100 token/s to 40.5 token/s. The dataset contains 288 valid observations with 100.0 percent coverage, meaning there are no missing-data limitations affecting this analysis.
Across the four-hour window, deepseek-v4-flash had the highest average output-token throughput at 99.44 token/s, while minimax-m3 was the weakest at 36.79 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which dropped from a peak of 116.52 token/s down to 4.27 token/s, reflecting a 62.9 percent coefficient of variation. Similarly, gemma4:31b exhibited a sharp downward trend, falling 46.3 percent to finish at 23.55 token/s. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage. Because the data is limited to this single rolling window and lacks any missing-data intervals, it cannot represent longer-term throughput baselines or diurnal patterns.
Across the four-hour window, deepseek-v4-flash is the strongest model with an average throughput of 103.13 token/s, while minimax-m3 is the weakest at 34.42 token/s. The most operationally significant volatility appears in gemma4:31b, which has a coefficient of variation of 51.2 percent and a downward trend of 53.3 percent, dropping from early highs near 99.66 token/s to a latest value of 21.14 token/s. By contrast, deepseek-v4-pro remains the most stable, with a coefficient of variation of 17.9 percent and an average of 85.77 token/s. The dataset has 100.0 percent coverage with 288 valid points, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, deepseek-v4-flash had the strongest average throughput at 101.66 token/s, while minimax-m3 was weakest at 35.12 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which ranged from 4.37 to 116.52 token/s with a 78.2 percent coefficient of variation, indicating severe instability despite a 182.8 percent upward trend. Additionally, gemma4:31b exhibited a sharp throughput decline of 31.8 percent, dropping from an early peak of 113.35 token/s to 22.92 token/s by the end. The dataset contains 288 valid observations across all models, achieving 100.0 percent coverage with no missing-data limitations.
Over the four-hour window, deepseek-v4-flash had the strongest average throughput at 107.22 token/s, while nemotron-3-ultra was weakest at 30.06 token/s. The most operationally significant volatility was nemotron-3-ultra’s extreme instability, with a coefficient of variation of 90.6 percent and a sharp upward trend of 319.6 percent, climbing from a low of 3.68 token/s to a peak of 116.52 token/s. In contrast, deepseek-v4-pro and glm-5.2 remained comparatively stable, averaging 83.72 and 90.15 token/s respectively. The dataset includes 288 valid observations across six models with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, deepseek-v4-flash achieved the strongest average throughput at 110.44 token/s, while nemotron-3-ultra was the weakest at 30.32 token/s. The most operationally significant volatility was nemotron-3-ultra's extreme instability, which had a coefficient of variation of 101.4 percent and a downward trend of 48.6 percent, dropping from an early peak of 124.36 token/s to a low of 3.68 token/s. In contrast, glm-5.2 maintained the most stable performance with a coefficient of variation of 27.2 percent. The dataset includes complete observations with 100.0 percent coverage and no missing-data limitations across all 288 valid points.
Across the four-hour window, deepseek-v4-flash is the strongest model with an average throughput of 109.74 token/s, while minimax-m3 is the weakest at 36.17 token/s. The most operationally significant trend is nemotron-3-ultra's severe throughput collapse, dropping 82.7 percent from an early peak of 159.16 token/s down to 15.03 token/s by the end. This model also exhibits extreme volatility with a coefficient of variation of 105.8 percent, including a low of 3.68 token/s. The dataset has complete coverage with 288 valid points and no missing-data limitation across all six models.
Over the four-hour window, deepseek-v4-flash had the highest average throughput at 107.68 token/s, while minimax-m3 was the weakest at 37.75 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which ranged from 3.68 to 159.16 token/s with a coefficient of variation of 86.1 percent and a downward trend of 40.2 percent, ending at 12.93 token/s. In contrast, deepseek-v4-flash trended upward by 24.2 percent, peaking at 158.03 token/s. The dataset records 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitation.
Across the four-hour window, deepseek-v4-flash had the highest average throughput at 97.68 token/s, while minimax-m3 was the weakest at 41.83 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which dropped from 152.86 token/s at 22:55 to 15.44 token/s at 23:00, reflecting its high coefficient of variation of 62.8 percent. Additionally, deepseek-v4-flash showed a strong upward trend, increasing by 33.8 percent over the period and reaching 127.69 token/s by 01:00. The dataset includes 288 valid observations across six models with 100.0 percent coverage, meaning there are no missing-data limitations to report for this interval.
Across the four-hour window, deepseek-v4-pro had the strongest average output-token throughput at 90.01 token/s, while minimax-m3 was the weakest at 43.78 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which ranged from 6.03 to 159.16 token/s with a coefficient of variation of 67.0 percent, indicating highly unstable generation speeds. In contrast, glm-5.2 maintained the most consistent performance, varying only between 21.76 and 115.46 token/s. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations. All models recorded exactly 48 five-minute samples, providing complete visibility into throughput trends during this period.
Over the four-hour window, glm-5.2 had the strongest average throughput at 86.14 token/s, while minimax-m3 was weakest at 43.15 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which ranged from 3.67 to 152.86 token/s with a coefficient of variation of 74.3 percent, including severe drops below 10 token/s near 19:55 and 22:40. In contrast, deepseek-v4-flash remained the most stable, maintaining an average of 84.73 token/s with a coefficient of variation of just 18.6 percent. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.
Over the four-hour window, deepseek-v4-flash had the strongest average output-token throughput at 83.99 token/s, while nemotron-3-ultra was weakest at 47.85 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which dropped to a minimum of 3.67 token/s and exhibited a coefficient of variation of 70.6 percent, alongside extreme swings such as a spike to 140.34 token/s. In contrast, deepseek-v4-flash remained stable with a coefficient of variation of 17.4 percent. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitation.
Across the four-hour window, deepseek-v4-flash achieved the highest average throughput at 85.98 token/s, while nemotron-3-ultra was the weakest at 39.88 token/s. The most operationally significant volatility came from nemotron-3-ultra, which exhibited a coefficient of variation of 69.8 percent and dropped to a minimum of 3.67 token/s. Additionally, gemma4:31b experienced a notable downward trend, falling 31.1 percent to a latest value of 46.23 token/s. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage. Because the data captures only this specific four-hour period, any missing-data limitation regarding broader daily or weekly throughput cycles cannot be assessed.
Across the four-hour window, deepseek-v4-flash had the strongest average throughput at 87.88 token/s, while minimax-m3 was weakest at 37.45 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which dropped sharply from 87.58 token/s at 19:10 to 3.67 token/s by 19:55, yielding a high coefficient of variation of 71.3 percent. Similarly, gemma4:31b exhibited extreme swings, ranging from 11.65 to 136.02 token/s. This analysis is constrained by a missing-data limitation: each model recorded 47 valid samples instead of the expected 48, leaving a small gap in the rolling five-minute observations.
Across the four-hour window, deepseek-v4-flash had the highest average output-token throughput at 87.14 token/s, while nemotron-3-ultra was the weakest at 43.02 token/s. The most operationally significant volatility appeared in gemma4:31b, which swung from a low of 9.86 token/s to a high of 136.02 token/s with a coefficient of variation of 66.8 percent, indicating highly unstable generation speeds. In contrast, deepseek-v4-flash maintained the most consistent performance, varying only between 34.94 token/s and 119.49 token/s. The dataset includes complete observations with 100.0 percent coverage and no missing-data limitations across all evaluated models.
Across the four-hour window, deepseek-v4-flash is the strongest model with an average throughput of 85.13 token/s, while minimax-m3 is the weakest at 32.0 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which has a coefficient of variation of 69.1 percent and throughput swings from a minimum of 3.41 token/s to a maximum of 116.34 token/s. glm-5.2 shows a notable downward trend of 25.4 percent, dropping from early highs above 110 token/s to 52.01 token/s by the end. The dataset includes 288 valid observations with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, deepseek-v4-flash had the strongest average output-token throughput at 81.56 token/s, while minimax-m3 was weakest at 32.28 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which ranged from 3.41 to 116.34 token/s with a coefficient of variation of 72.2 percent, indicating highly unstable generation speeds. Additionally, glm-5.2 exhibited a sharp end-of-window throughput drop, falling from 113.07 token/s at 15:50 to just 12.38 token/s by 17:00. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 86.86 token/s, while nemotron-3-ultra was the weakest at 32.58 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which fluctuated wildly between 3.41 and 76.87 token/s with a coefficient of variation of 69.6 percent, indicating highly unstable generation speeds. Conversely, deepseek-v4-flash showed steady performance, trending upward by 2.9 percent to average 78.71 token/s. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 87.36 token/s, while nemotron-3-ultra was the weakest at 32.2 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which exhibited a 75.1 percent coefficient of variation, swinging between a maximum of 95.51 token/s and a minimum of 3.41 token/s, alongside a 31.3 percent downward trend. In contrast, glm-5.2 maintained comparative stability with a 20.3 percent coefficient of variation. The dataset contains 288 valid points, achieving 100 percent coverage across all 48 expected samples per model, meaning there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 87.39 token/s, while minimax-m3 was the weakest at 38.42 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which exhibited a 71.6% coefficient of variation, frequent sharp drops to near 5 token/s, and a -24.5% downward trend. In contrast, deepseek-v4-pro showed steadier performance, trending upward by 4.6% with a 27.7% coefficient of variation. The dataset contains 288 valid observations across six models, achieving 100.0% coverage with no missing-data limitations.
Across the four-hour window, glm-5.2 is the strongest model by average throughput at 84.26 token/s, narrowly ahead of deepseek-v4-flash at 84.01 token/s. The weakest is minimax-m3 at 37.91 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which has a coefficient of variation of 71.7 percent and a range from 5.19 to 131.08 token/s, indicating highly unstable output rates. Dataset coverage is complete with 288 valid points across all six models, so there are no missing-data limitations affecting this review.
Across the four-hour window, glm-5.2 delivered the highest average output-token throughput at 81.62 token/s, while minimax-m3 was the weakest at 34.79 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which ranged from 8.59 to 139.62 token/s with a coefficient of variation of 69.9 percent, indicating severe throughput instability compared to glm-5.2's 25.6 percent. Deepseek-v4-flash also showed notable swings, dropping to 17.15 token/s at 08:10 before peaking at 118.94 token/s at 09:10. The dataset contains 288 valid observations across six models with 100.0 percent coverage, so no missing-data limitation affects this specific interval.
Across the four-hour window, deepseek-v4-flash had the strongest average throughput at 82.06 token/s, while minimax-m3 was the weakest at 31.45 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which ranged from 6.9 to 139.62 token/s with a coefficient of variation of 69.6 percent, indicating highly unstable output generation. In contrast, glm-5.2 maintained a steadier average of 78.13 token/s with a lower coefficient of variation of 26.3 percent. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations, providing a complete view of system performance.
Across the four-hour window, deepseek-v4-flash achieved the strongest average throughput at 82.24 token/s, while minimax-m3 was the weakest at 30.54 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which fluctuated wildly between 6.9 and 139.62 token/s, yielding a coefficient of variation of 69.3 percent. Additionally, gemma4:31b exhibited a sharp downward trend, dropping 34.4 percent to a late low of 20.23 token/s. The dataset contains 288 valid observations across six models, achieving exactly 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 82.72 token/s, while minimax-m3 was the weakest at 33.81 token/s. The most operationally significant trend is deepseek-v4-flash improving by 41.7 percent to 119.21 token/s, contrasting with nemotron-3-ultra, which declined 31.9 percent to 58.97 token/s. High volatility plagues nemotron-3-ultra, evidenced by its 60.6 percent coefficient of variation and a maximum drop to 6.9 token/s. The dataset contains 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 88.37 token/s, while minimax-m3 is the weakest at 39.32 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits extreme throughput swings, dropping to a minimum of 5.45 token/s before peaking at 152.72 token/s, yielding a coefficient of variation of 57.6 percent. This high variance contrasts with the steadier glm-5.2, which maintains a lower coefficient of variation of 23.2 percent. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.