Across the four-hour window, gemma4:31b had the strongest average output-token throughput at 95.83 token/s, while nemotron-3-ultra was weakest at 36.9 token/s. The most operationally significant volatility appeared in deepseek-v4-flash, which dropped sharply from 95.72 token/s at 19:55 to 12.87 token/s at 20:35 before recovering to 93.53 token/s by 20:55. This instability reflects its high coefficient of variation of 41.6 percent. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitation.
Hourly performance insights
Summaries of rolling four-hour performance data
Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 87.74 token/s, while nemotron-3-ultra was the weakest at 35.81 token/s. The most operationally significant volatility appeared in gemma4:31b, which jumped from a 16:20 low of 5.82 token/s to a 17:25 high of 143.24 token/s, yielding a 63.9 percent coefficient of variation. Deepseek-v4-pro also showed sharp instability, dropping to 2.92 token/s at 17:45. The dataset contains 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitation.
Across the four-hour window, glm-5.2 is the strongest model by average throughput at 82.3 token/s, while nemotron-3-ultra is the weakest at 32.37 token/s. The most operationally significant volatility comes from gemma4:31b, which exhibits a massive 478.5 percent upward trend, jumping from a low of 5.82 token/s early in the window to a maximum of 153.29 token/s. In contrast, minimax-m3 remains highly stable, averaging 45.84 token/s with a low coefficient of variation of 19.8 percent. The dataset includes 288 valid observations, achieving 100.0 percent coverage across all six models, meaning there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 72.48 token/s, while gemma4:31b was the weakest at 30.02 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which dropped sharply from 92.29 token/s at 15:25 to 4.23 token/s at 15:35. Similarly, deepseek-v4-pro experienced a severe single-interval drop to 2.92 token/s at 17:45. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations. This complete dataset allows reliable observation of extreme throughput fluctuations, including gemma4:31b surging to 143.24 token/s at 17:25.
Across the four-hour window, deepseek-v4-pro had the strongest average output-token throughput at 74.36 token/s, while gemma4:31b was the weakest at 13.26 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which dropped from 115.84 token/s at 13:15 to 1.93 token/s at 14:20, reflecting a coefficient of variation of 85.7 percent. Similarly, deepseek-v4-flash experienced severe instability, plummeting from 101.82 token/s at 14:15 to 4.23 token/s at 15:35. The dataset includes 288 valid observations, achieving 100.0 percent coverage across all six models, meaning there are no missing-data limitations affecting this analysis.
Over the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 73.1 token/s, while gemma4:31b was the weakest at 14.64 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which fluctuated from a low of 1.93 token/s to a high of 115.84 token/s with a coefficient of variation of 91.7 percent, indicating severe instability. Deepseek-v4-flash also showed notable volatility, dropping to 4.23 token/s near 15:35. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 73.08 token/s, while gemma4:31b was the weakest at 23.64 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which dropped to 1.93 token/s despite a 30.51 token/s average, and gemma4:31b, which fell 68.3 percent from an initial 114.14 token/s to a low of 7.2 token/s. In contrast, minimax-m3 maintained the most stable performance, averaging 44.26 token/s with a 22.4 percent coefficient of variation. The dataset includes 288 valid observations across six models with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 is the strongest model by average throughput at 77.79 token/s, while nemotron-3-ultra is the weakest at 32.87 token/s. The most operationally significant volatility comes from gemma4:31b, which has a coefficient of variation of 85.8% and a downward trend of -77.5%, dropping from a peak of 140.22 token/s to a latest reading of 7.2 token/s. Nemotron-3-ultra also shows severe instability, bottoming out at 2.24 token/s. In contrast, minimax-m3 is the most stable model, averaging 42.32 token/s with a coefficient of variation of just 20.6%. The dataset has 100.0% coverage with 288 valid points, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the highest average throughput at 74.9 token/s, while nemotron-3-ultra was the weakest at 27.15 token/s. Operationally, gemma4:31b exhibited the most significant volatility and downward trend, dropping from a peak of 140.22 token/s to a low of 9.43 token/s, reflecting a 46.6 percent trend decrease. Nemotron-3-ultra also showed extreme instability with a 95.7 percent coefficient of variation, frequently falling below 3 token/s near 12:25. The dataset contains 288 valid observations across all six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, deepseek-v4-pro had the highest average output-token throughput at 75.45 token/s, while nemotron-3-ultra was the weakest at 29.67 token/s. The most operationally significant volatility came from nemotron-3-ultra, which swung between 1.82 and 108.64 token/s with a coefficient of variation of 98.6 percent, indicating highly unstable generation speeds. In contrast, minimax-m3 offered the most stable performance, averaging 43.17 token/s with a coefficient of variation of just 17.1 percent. The dataset includes complete observations with 100.0 percent coverage and no missing-data limitations across all 288 valid points, providing a reliable basis for this review.
Across the four-hour window, deepseek-v4-pro had the strongest average output-token throughput at 76.2 token/s, while nemotron-3-ultra was weakest at 25.28 token/s. The most operationally significant volatility came from nemotron-3-ultra, which ranged from 1.82 to 96.1 token/s with a coefficient of variation of 99.8 percent, including a sustained drop to near-zero rates between 08:40 and 09:10. deepseek-v4-flash also showed instability, dipping to 9.02 token/s. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitation.
Across the four-hour window, deepseek-v4-pro delivered the strongest average throughput at 74.58 token/s, while nemotron-3-ultra was the weakest at 21.6 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which dropped from a peak of 93.93 token/s to sustained lows near 1.82 token/s between 08:40 and 09:15 UTC, yielding a coefficient of variation of 100.7%. In contrast, minimax-m3 maintained the most stable performance, averaging 41.96 token/s with a standard deviation of just 6.95 token/s. The dataset contains 288 valid observations with 100.0% coverage, meaning there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 88.54 token/s, while nemotron-3-ultra was the weakest at 27.03 token/s. The most operationally significant volatility came from nemotron-3-ultra, which exhibited extreme instability with a coefficient of variation of 93.0 percent and sharp drops to a minimum of 2.0 token/s. Additionally, all six models experienced negative throughput trends over the observed period, with nemotron-3-ultra declining by 32.5 percent and gemma4:31b by 25.9 percent. The dataset contains 288 valid observations, achieving 100.0 percent coverage across all models, meaning there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 had the strongest average output-token throughput at 91.61 token/s, while nemotron-3-ultra was weakest at 33.86 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which dropped from an early peak of 109.19 token/s to a sustained low near 2.74 token/s between 05:25 and 05:50, reflecting an 87.8 percent coefficient of variation and a negative trend of 64.6 percent. gemma4:31b also showed notable instability, swinging between 27.19 and 155.32 token/s. The dataset contains 288 valid points with 100.0 percent coverage, so no missing-data limitation affects this specific analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 96.13 token/s, while nemotron-3-ultra was weakest at 31.59 token/s. The most operationally significant volatility belongs to nemotron-3-ultra, which swung from a low of 1.9 token/s to a high of 109.19 token/s, yielding a coefficient of variation of 96.7 percent. This extreme instability contrasts with minimax-m3, which maintained steadier performance around its 45.11 token/s average. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations. All models reported complete sample counts, providing a fully reliable view of output-token throughput for this period.
Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 95.47 token/s, while nemotron-3-ultra was the weakest at 38.14 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which exhibited extreme instability with a coefficient of variation of 77.2 percent, dropping as low as 1.9 token/s before spiking to 109.19 token/s. In contrast, minimax-m3 maintained the most stable performance, averaging 45.97 token/s with a standard deviation of only 8.2 token/s. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 98.03 token/s, while nemotron-3-ultra is the weakest at 36.01 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which has a coefficient of variation of 78.6 percent and drops to a minimum of 1.9 token/s. Conversely, deepseek-v4-flash shows a notable upward trend of 37.8 percent, rising from early lows to a latest throughput of 59.56 token/s. The dataset includes 288 valid points across six models, achieving 100.0 percent coverage. Because the data lacks concurrent request counts, throughput drops cannot be definitively attributed to capacity limits.
Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 100.04 token/s, while nemotron-3-ultra is the weakest at 30.33 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which has a coefficient of variation of 82.2 percent and drops to a minimum of 1.9 token/s. Deepseek-v4-pro also exhibits severe instability, plunging to 0.61 token/s at 00:25 despite an average of 81.7 token/s. The dataset contains 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, glm-5.2 had the strongest average output-token throughput at 96.0 token/s, while nemotron-3-ultra was weakest at 36.58 token/s. The most operationally significant volatility appeared in deepseek-v4-pro and deepseek-v4-flash, which experienced severe throughput drops to 0.61 token/s and 8.2 token/s respectively, indicating intermittent processing stalls. Additionally, nemotron-3-ultra exhibited extreme instability with a coefficient of variation of 72.2 percent. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations. This complete sample set confirms all models maintained continuous reporting throughout the entire monitoring period.
Across the four-hour window, glm-5.2 delivered the highest average throughput at 98.22 token/s, while nemotron-3-ultra was the weakest at 34.73 token/s. The most operationally significant volatility occurred in deepseek-v4-pro, which experienced a severe throughput drop to 0.61 token/s at 00:25 despite an average of 84.16 token/s. Similarly, glm-5.2 briefly plummeted to 12.74 token/s at 23:45. Although the dataset reports 100.0 percent coverage across all 288 valid observations, this analysis is limited by the absence of missing-data context, meaning these extreme throughput dips cannot be attributed to specific outages or collection gaps.
Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 88.49 token/s, while nemotron-3-ultra was the weakest at 37.46 token/s. The most operationally significant volatility appeared in deepseek-v4-flash and nemotron-3-ultra, which exhibited extreme swings; nemotron-3-ultra fluctuated between a low of 5.5 token/s and a high of 118.14 token/s, yielding a coefficient of variation of 77.2 percent. In contrast, minimax-m3 maintained the most stable performance, averaging 53.43 token/s with a standard deviation of only 9.06 token/s. The dataset contains 288 valid observations, achieving 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 88.24 token/s, while nemotron-3-ultra was the weakest at 35.65 token/s. The most operationally significant volatility appears in deepseek-v4-flash and nemotron-3-ultra, which exhibited extreme swings; nemotron-3-ultra fluctuated between 5.67 and 118.14 token/s, and deepseek-v4-flash dropped to 2.77 token/s. In contrast, minimax-m3 maintained the most stable performance, averaging 53.25 token/s with a low coefficient of variation of 14.9 percent. The dataset contains 288 valid points across all six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 81.8 token/s, while gemma4:31b was the weakest at 33.38 token/s. The most operationally significant volatility appeared in deepseek-v4-flash, which exhibited extreme swings between 2.77 and 112.11 token/s and a negative trend of 30.7 percent, contrasting sharply with the stable minimax-m3 at 53.36 token/s average and a 14.8 percent coefficient of variation. Although overall dataset coverage is 100.0 percent across 288 valid points, the five-minute aggregation interval masks sub-interval latency spikes, limiting visibility into brief throughput drops.
Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 80.82 token/s, while nemotron-3-ultra was the weakest at 28.5 token/s. The most operationally significant volatility appeared in gemma4:31b, which spiked from 18.13 token/s at 18:50 to 115.42 token/s at 18:40, contributing to a high coefficient of variation of 60.9 percent. In contrast, minimax-m3 maintained the most stable performance, averaging 51.19 token/s with a standard deviation of just 9.28 token/s. The dataset includes 288 valid observations, achieving 100.0 percent coverage across all six models, meaning there are no missing-data limitations affecting this specific analysis.
Across the four-hour observation window, glm-5.2 delivered the strongest average output-token throughput at 75.15 token/s, while gemma4:31b was the weakest at 24.55 token/s. The most operationally significant volatility appeared in deepseek-v4-flash, which exhibited extreme swings between 9.81 and 104.02 token/s, reflecting a coefficient of variation of 50.6 percent. Similarly, gemma4:31b showed high instability, plunging to 9.29 token/s before spiking to 115.42 token/s. In contrast, minimax-m3 maintained the most stable performance, fluctuating narrowly between 28.61 and 64.12 token/s. The dataset contains 288 valid observations across six models, achieving complete coverage with no missing-data limitations affecting the throughput analysis.
Across the four-hour window, glm-5.2 delivered the highest average output-token throughput at 71.78 token/s, while gemma4:31b was the weakest at 15.05 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which dropped from a peak of 111.71 token/s to a low of 4.05 token/s, reflecting a 77.9 percent coefficient of variation and a -53.7 percent trend. Similarly, nemotron-3-ultra showed extreme instability with a 79.5 percent coefficient of variation. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.
Over the four-hour window, glm-5.2 had the strongest average output-token throughput at 76.35 token/s, while gemma4:31b was the weakest at 15.52 token/s. The most operationally significant trend was deepseek-v4-flash declining by 44.3 percent, dropping from an early peak of 111.71 token/s to a latest reading of 4.05 token/s, alongside high volatility with a coefficient of variation of 71.9 percent. Nemotron-3-ultra also showed severe volatility, swinging between 3.62 and 85.73 token/s. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitation.
Across the four-hour window, glm-5.2 is the strongest model by average throughput at 80.3 token/s, narrowly beating deepseek-v4-pro at 80.22 token/s. The weakest is gemma4:31b at 15.9 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which has a coefficient of variation of 70.1 percent and throughput swings from 3.62 to 86.45 token/s. Deepseek-v4-flash also shows high instability, with a 51.8 percent coefficient of variation and a drop to 9.37 token/s. The dataset records 288 valid points across six models with 100.0 percent coverage, so there are no missing-data limitations affecting this specific interval.
Across the four-hour window, deepseek-v4-pro had the strongest average output-token throughput at 84.19 token/s, while gemma4:31b was the weakest at 14.1 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which dropped from a peak of 93.25 token/s to a low of 3.62 token/s with a coefficient of variation of 68.2 percent, alongside a negative trend of 24.9 percent. deepseek-v4-flash also showed high instability, ranging from 9.37 to 113.57 token/s. The dataset contains complete coverage with 288 valid points and no missing-data limitation.
Across the four-hour window, deepseek-v4-pro had the strongest average output-token throughput at 84.69 token/s, while gemma4:31b was the weakest at 16.22 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which declined by 10.3 percent overall and fluctuated heavily between 3.62 and 102.0 token/s with a coefficient of variation of 65.2 percent. deepseek-v4-flash also showed instability, dropping to 9.47 token/s despite averaging 57.82 token/s. The dataset includes 288 valid observations across six models with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, deepseek-v4-pro had the strongest average throughput at 84.24 token/s, while gemma4:31b was the weakest at 16.89 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which swung between 5.92 and 102.0 token/s with a coefficient of variation of 67.3 percent, and deepseek-v4-flash, which ranged from 11.27 to 113.57 token/s. Such extreme swings indicate highly unstable output generation. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage. While no missing-data limitation exists for the expected samples, the dataset lacks concurrent request counts, preventing any analysis of how load influenced these throughput variations.
Across the four-hour window, deepseek-v4-pro delivered the strongest average output-token throughput at 83.75 token/s, while gemma4:31b was the weakest at 17.45 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which ranged from 4.83 to 102.0 token/s with a 69.3 percent coefficient of variation, indicating highly unstable generation speeds despite a positive 57.7 percent trend. Deepseek-v4-flash also showed notable instability, dropping to 11.27 token/s. The dataset contains 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitations. All models recorded exactly 48 samples, matching the expected count perfectly.
Over the four-hour window, glm-5.2 delivered the strongest average throughput at 88.6 token/s, while gemma4:31b was the weakest at 19.45 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which varied wildly between 4.83 and 102.0 token/s with a coefficient of variation of 66.2 percent. Deepseek-v4-flash also showed extreme instability, dropping as low as 7.16 token/s. Most models exhibited negative trend percentages, indicating declining throughput toward the end of the period. The dataset contains 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitation.
Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 90.41 token/s, while gemma4:31b was the weakest at 20.45 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which dropped 48.7 percent over the period and hit a low of 4.83 token/s, alongside a coefficient of variation of 64.6 percent. deepseek-v4-flash also showed instability, swinging between 7.16 and 103.08 token/s. The dataset records 288 valid points across six models with 100.0 percent coverage, so no missing-data limitation affects this specific interval.
Across the four-hour window, glm-5.2 delivered the highest average output-token throughput at 93.81 token/s, while gemma4:31b was the weakest at 28.48 token/s. The most operationally significant volatility appeared in gemma4:31b, which spiked to 144.09 token/s before trending down by 54.0 percent to a latest reading of 15.77 token/s. In contrast, minimax-m3 maintained the most stable performance, averaging 44.6 token/s with a low coefficient of variation of 18.8 percent. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage. Because the dataset lacks concurrent request concurrency counts, throughput drops cannot be definitively attributed to queueing or capacity limits.
Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 92.06 token/s, while gemma4:31b was the weakest at 39.82 token/s. The most operationally significant volatility appeared in gemma4:31b, which dropped 61.5 percent overall, falling from early peaks near 144.09 token/s to a latest reading of 16.55 token/s. In contrast, minimax-m3 remained stable with a coefficient of variation of 21.3 percent and an average of 47.4 token/s. The dataset includes 288 valid observations, achieving 100.0 percent coverage across all models, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the highest average output-token throughput at 90.71 token/s, while gemma4:31b was the weakest at 41.22 token/s. Operationally, throughput volatility was highly significant, with gemma4:31b and nemotron-3-ultra exhibiting extreme swings; nemotron-3-ultra fluctuated between a low of 8.47 token/s and a high of 98.43 token/s, yielding a coefficient of variation of 57.0 percent. Deepseek-v4-pro showed the most positive operational trend, increasing by 10.8 percent to a latest throughput of 99.99 token/s. The dataset contains 288 valid observations with 100.0 percent coverage, meaning there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 93.27 token/s, while minimax-m3 is the weakest at 47.42 token/s. The most operationally significant volatility appears in gemma4:31b, which has a coefficient of variation of 72.4 percent and throughput swings from a minimum of 8.86 token/s to a maximum of 144.09 token/s. Deepseek-v4-flash also shows notable instability, dropping to 11.03 token/s before reaching 129.57 token/s. The dataset includes 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 97.3 token/s, while minimax-m3 is the weakest at 50.64 token/s. The most operationally significant volatility appears in gemma4:31b, which has a coefficient of variation of 71.9 percent and a negative trend of 32.1 percent, dropping from a peak of 152.72 token/s to a low of 8.86 token/s. Deepseek-v4-flash and nemotron-3-ultra also show high volatility, with coefficients of variation of 48.3 percent and 55.1 percent respectively. The dataset includes 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 108.85 token/s, while minimax-m3 is the weakest at 53.27 token/s. The most operationally significant volatility comes from gemma4:31b, which has a coefficient of variation of 50.1 percent and a negative trend of 28.0 percent, dropping from a peak of 159.57 token/s to a low of 17.18 token/s late in the period. deepseek-v4-flash also shows severe instability, ranging from 9.94 to 119.26 token/s. The dataset includes 288 valid points across six models with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 111.46 token/s, while minimax-m3 is the weakest at 53.11 token/s. The most operationally significant volatility appears in deepseek-v4-flash and nemotron-3-ultra, which exhibit extreme throughput swings; nemotron-3-ultra drops to a minimum of 3.98 token/s, and deepseek-v4-flash falls to 9.94 token/s. This volatility indicates highly unstable output generation for these specific models. The dataset contains 288 valid observations across all six models, achieving 100.0 percent coverage. Because the dataset lacks any missing-data limitations, the observed throughput patterns fully represent the complete four-hour monitoring period without gaps.
Across the four-hour observation window, glm-5.2 is the strongest model with an average throughput of 110.8 token/s, while minimax-m3 is the weakest at 50.49 token/s. The most operationally significant volatility appears in nemotron-3-ultra and deepseek-v4-flash, which exhibit severe throughput instability. Nemotron-3-ultra fluctuates wildly between 3.98 and 134.94 token/s with a coefficient of variation of 56.0 percent, and deepseek-v4-flash drops to lows of 11.47 token/s despite reaching 113.04 token/s. In contrast, glm-5.2 maintains the most stable performance. The dataset contains 288 valid observations with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 107.51 token/s, while minimax-m3 is the weakest at 48.67 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits extreme swings between 3.98 and 134.94 token/s, yielding a coefficient of variation of 66.6 percent despite an upward trend of 71.9 percent. In contrast, glm-5.2 maintains the most stable performance with a coefficient of variation of just 14.8 percent. The dataset contains 288 valid observations across all six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 105.67 token/s, while nemotron-3-ultra is the weakest at 41.8 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which has a coefficient of variation of 75.1 percent and swings between 3.98 and 106.45 token/s. deepseek-v4-flash also shows instability, dropping to 6.13 token/s. In contrast, glm-5.2 maintains the most stable performance with a coefficient of variation of just 14.7 percent. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 103.62 token/s, while nemotron-3-ultra is the weakest at 39.17 token/s. The most operationally significant volatility comes from nemotron-3-ultra and deepseek-v4-flash, which exhibit extreme throughput swings; nemotron-3-ultra fluctuates between 7.08 and 106.45 token/s with a coefficient of variation of 68.9 percent, and deepseek-v4-flash ranges from 6.13 to 119.9 token/s. In contrast, glm-5.2 maintains the most stable performance, varying only from 63.75 to 138.13 token/s. The dataset contains no missing-data limitation, as all six models report a complete 100.0 percent coverage across 288 valid observations.
Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 97.11 token/s, while nemotron-3-ultra is the weakest at 30.36 token/s. The most operationally significant volatility comes from gemma4:31b, which fluctuated heavily between a minimum of 6.18 token/s and a maximum of 135.6 token/s, yielding a coefficient of variation of 72.5 percent. Similarly, deepseek-v4-flash exhibited extreme instability, dropping to 6.13 token/s. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 95.74 token/s, while nemotron-3-ultra was weakest at 27.42 token/s. The most operationally significant volatility appeared in gemma4:31b, which dropped sharply from 112.14 token/s at 13:05 to a low of 6.18 token/s at 15:50, yielding a coefficient of variation of 67.5 percent. In contrast, minimax-m3 remained stable with an average of 43.65 token/s and a coefficient of variation of 20.8 percent. The dataset includes 288 valid observations across six models with 100.0 percent coverage, so no missing-data limitation affects this specific review.
Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 96.58 token/s, while nemotron-3-ultra was the weakest at 22.45 token/s. The most operationally significant volatility appeared in gemma4:31b, which ranged from 5.89 to 145.23 token/s with a 79.7 percent coefficient of variation and a -58.8 percent trend, indicating severe throughput degradation over time. In contrast, minimax-m3 maintained the most stable performance, averaging 45.12 token/s with a 19.2 percent coefficient of variation. Although the dataset reports 100.0 percent coverage across all 288 valid points, this analysis is limited by the absence of concurrent request volume data, preventing correlation of throughput drops with load spikes.
Across the four-hour window, glm-5.2 is the strongest model by average throughput at 102.18 token/s, while nemotron-3-ultra is the weakest at 22.67 token/s. The most operationally significant volatility appears in gemma4:31b, which has a coefficient of variation of 70.6% and throughput swings from a minimum of 5.89 token/s to a maximum of 145.23 token/s. Deepseek-v4-flash also shows notable instability, dropping to 6.18 token/s at 14:35 despite a 60.63 token/s average. The dataset records 288 valid points across six models, achieving 100.0% coverage with no missing-data limitation.
Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 105.85 token/s, while nemotron-3-ultra was the weakest at 26.57 token/s. The most operationally significant volatility appeared in gemma4:31b, which fluctuated heavily between a high of 171.29 token/s and a low of 5.89 token/s, yielding a coefficient of variation of 66.0 percent. In contrast, minimax-m3 maintained the most stable performance, averaging 43.85 token/s with a 16.5 percent coefficient of variation. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.