← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1205 retained summaries
4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 105.85 token/s, while nemotron-3-ultra was the weakest at 26.57 token/s. The most operationally significant volatility appeared in gemma4:31b, which fluctuated heavily between a high of 171.29 token/s and a low of 5.89 token/s, yielding a coefficient of variation of 66.0 percent. In contrast, minimax-m3 maintained the most stable performance, averaging 43.85 token/s with a 16.5 percent coefficient of variation. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 108.03 token/s, while nemotron-3-ultra was the weakest at 27.18 token/s. The most operationally significant volatility appeared in gemma4:31b, which dropped sharply from a peak of 171.29 token/s down to 5.89 token/s before partially recovering, reflecting a coefficient of variation of 61.2 percent. In contrast, minimax-m3 maintained the most stable performance, averaging 45.43 token/s with a low coefficient of variation of 15.1 percent. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour observation window, glm-5.2 delivered the strongest average output-token throughput at 110.35 token/s, while nemotron-3-ultra was the weakest at 29.03 token/s. The most operationally significant volatility occurred in gemma4:31b, which experienced a severe throughput decline starting around 11:00 UTC, dropping from roughly 107.54 token/s at 10:55 to a low of 16.03 token/s by 12:00. This sharp degradation contrasts with its earlier peak of 171.29 token/s at 10:25. The dataset contains complete telemetry with 100.0 percent coverage across all 288 valid observations, meaning there are no missing-data limitations affecting this specific performance review.

4-hour window · 288 points · 100.0% coverage

Over the four-hour window, glm-5.2 had the strongest average throughput at 110.81 token/s, while nemotron-3-ultra was weakest at 32.23 token/s. The most operationally significant volatility came from nemotron-3-ultra, which ranged from 3.14 to 97.92 token/s with a coefficient of variation of 78.9 percent, and deepseek-v4-flash, which dropped to 6.49 token/s despite an average of 73.34 token/s. Deepseek-v4-pro showed the most stable performance, varying only between 56.03 and 120.8 token/s. The dataset includes 288 valid observations across six models with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 118.32 token/s, while nemotron-3-ultra is the weakest at 34.42 token/s. The most operationally significant volatility comes from nemotron-3-ultra and deepseek-v4-flash, which exhibit extreme throughput swings; nemotron-3-ultra fluctuates between 3.14 and 116.65 token/s with a coefficient of variation of 88.1 percent, and deepseek-v4-flash ranges from 6.49 to 126.21 token/s. In contrast, deepseek-v4-pro is the most stable, maintaining 95.47 token/s on average with a low coefficient of variation of 14.2 percent. The dataset contains 288 valid observations with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 117.74 token/s, while nemotron-3-ultra was the weakest at 40.42 token/s. The most operationally significant volatility appeared in deepseek-v4-flash and nemotron-3-ultra, which exhibited severe throughput swings; nemotron-3-ultra dropped to a minimum of 3.14 token/s and deepseek-v4-flash to 6.49 token/s, indicating highly unstable generation rates. In contrast, deepseek-v4-pro maintained the most stable performance with a coefficient of variation of 13.5 percent. The dataset contains 288 valid observations across all six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Over the four-hour window, glm-5.2 had the strongest average output-token throughput at 116.92 token/s, while nemotron-3-ultra was weakest at 43.44 token/s. The most operationally significant volatility came from nemotron-3-ultra, which had a coefficient of variation of 79.5% and a downward trend of -20.1%, dropping to a minimum of 1.95 token/s. Deepseek-v4-flash also showed severe instability, with a coefficient of variation of 48.0% and a minimum of 8.75 token/s. In contrast, deepseek-v4-pro was the most stable at 94.39 token/s with a coefficient of variation of 13.9%. The dataset includes 288 valid points across six models with 100.0% coverage, so there are no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model by average throughput at 119.58 token/s, while nemotron-3-ultra is the weakest at 40.43 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which has a coefficient of variation of 85.3% and swings from a high of 127.3 token/s down to 1.95 token/s. Similarly, deepseek-v4-flash shows high instability with a 45.8% coefficient of variation and a low of 9.47 token/s. In contrast, deepseek-v4-pro is the most stable at 13.3%. The dataset includes 288 valid observations with 100.0% coverage, so there are no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 had the strongest average output-token throughput at 115.71 token/s, while nemotron-3-ultra was weakest at 44.78 token/s. The most operationally significant volatility came from nemotron-3-ultra, which had a coefficient of variation of 81.3 percent and dropped to a minimum of 1.95 token/s. Deepseek-v4-flash also showed instability, falling to 9.47 token/s. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 114.91 token/s, while minimax-m3 is the weakest at 46.8 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which has a coefficient of variation of 80.5 percent and a downward trend of 40.7 percent, dropping from a maximum of 132.61 token/s to a minimum of 1.95 token/s. Deepseek-v4-flash also shows high volatility with a coefficient of variation of 42.4 percent. The dataset has 100.0 percent coverage with 288 valid points, so there are no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 114.12 token/s, while minimax-m3 is the weakest at 47.94 token/s. The most operationally significant volatility appears in deepseek-v4-flash and nemotron-3-ultra. Deepseek-v4-flash exhibits extreme swings, dropping to a minimum of 7.41 token/s despite a 71.47 token/s average. Nemotron-3-ultra shows a severe downward trend, declining 27.1 percent to a latest throughput of just 3.73 token/s. In contrast, deepseek-v4-pro remains the most stable, maintaining a steady 99.28 token/s average with a low coefficient of variation of 12.4 percent. The dataset contains no missing-data limitation, as all models achieved exactly 48 valid samples for 100.0 percent coverage.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 114.97 token/s, while minimax-m3 is the weakest at 49.38 token/s. The most operationally significant volatility appears in deepseek-v4-flash and nemotron-3-ultra, which exhibit extreme throughput swings; deepseek-v4-flash drops to a minimum of 7.41 token/s with a coefficient of variation of 51.6%, and nemotron-3-ultra falls to 4.82 token/s with a negative trend of -13.6%. The dataset contains 288 valid observations across all six models, achieving 100% coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 109.94 token/s, while minimax-m3 is the weakest at 48.82 token/s. The most operationally significant volatility appears in deepseek-v4-flash and nemotron-3-ultra, which exhibit extreme throughput swings. Deepseek-v4-flash dropped to a minimum of 7.41 token/s and nemotron-3-ultra to 10.11 token/s, with both models showing coefficient of variation values near 49 percent. Dataset coverage is complete at 100 percent across all 288 valid observations, meaning there are no missing-data limitations impacting this specific rolling window analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model by average throughput at 105.7 token/s, narrowly exceeding gemma4:31b at 105.48 token/s. The weakest model is minimax-m3, averaging 49.57 token/s. Operationally, deepseek-v4-flash and nemotron-3-ultra exhibit extreme volatility, with coefficients of variation of 51.1 percent and 50.7 percent respectively. Deepseek-v4-flash throughput frequently collapses, dropping to a minimum of 7.41 token/s, while nemotron-3-ultra shows a steep upward trend of 26.5 percent over the period. The dataset contains 288 valid observations, achieving 100.0 percent coverage with no missing-data limitations across the expected 48 samples per model.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 104.55 token/s, while minimax-m3 was the weakest at 50.35 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which swung between 10.11 and 119.01 token/s with a coefficient of variation of 56.1 percent, indicating highly unstable generation speeds. Deepseek-v4-flash also showed severe instability, dropping to a minimum of 8.24 token/s despite an average of 77.24 token/s. The dataset contains 288 valid observations across all six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 109.59 token/s, while nemotron-3-ultra is the weakest at 47.69 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which has a coefficient of variation of 68.6% and throughput swings between 7.52 and 113.6 token/s. Similarly, deepseek-v4-flash and deepseek-v4-pro experience steep drops to 7.34 and 7.9 token/s respectively. The dataset has 100.0% coverage with 288 valid points, meaning there are no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 109.46 token/s, while nemotron-3-ultra was the weakest at 46.94 token/s. The most operationally significant volatility came from nemotron-3-ultra, which fluctuated wildly between 7.52 and 119.76 token/s with a coefficient of variation of 71.0 percent, and deepseek-v4-flash, which dropped to 3.38 token/s. Despite this volatility, nemotron-3-ultra showed a positive trend of 31.2 percent over the period. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 100.68 token/s, while nemotron-3-ultra was the weakest at 47.73 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which dropped from 111.16 token/s at 17:25 to 8.09 token/s at 17:55, reflecting its high coefficient of variation of 69.2 percent. Additionally, deepseek-v4-pro exhibited a sharp downward trend, falling 16.3 percent over the period to a late low of 7.9 token/s at 20:00. Although current coverage is complete, any missing-data limitation would obscure these severe minute-level fluctuations and distort true capacity planning.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 103.45 token/s, while nemotron-3-ultra was the weakest at 47.32 token/s. The most operationally significant volatility appears in deepseek-v4-flash and nemotron-3-ultra, which exhibited extreme throughput swings; flash ranged from 3.38 to 127.9 token/s with a 53.1% coefficient of variation, and nemotron ranged from 7.52 to 119.76 token/s with a 65.0% coefficient of variation. Additionally, nemotron-3-ultra showed a sharp downward trend, dropping 24.9% over the period to a latest throughput of 7.52 token/s. The dataset contains no missing-data limitation, as all six models achieved 100.0% coverage across 288 valid observations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 101.39 token/s, while minimax-m3 was the weakest at 46.33 token/s. The most operationally significant volatility occurred in deepseek-v4-flash and nemotron-3-ultra, which exhibited extreme throughput swings despite their lower averages. Deepseek-v4-flash dropped to a minimum of 2.68 token/s with a coefficient of variation of 60.1%, and nemotron-3-ultra reached a low of 8.09 token/s with a coefficient of variation of 63.3%. The dataset contains 288 valid observations across all six models, achieving 100.0% coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the highest average throughput at 102.01 token/s, while nemotron-3-ultra was the weakest at 43.29 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which exhibited extreme swings, plummeting to a minimum of 2.68 token/s and recording a coefficient of variation of 56.8 percent. Similarly, glm-5.2 suffered a severe throughput drop to 5.9 token/s near 17:25 before recovering. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 112.45 token/s, while nemotron-3-ultra is the weakest at 32.2 token/s. The most operationally significant volatility appears in deepseek-v4-flash and nemotron-3-ultra, which exhibit coefficient of variation values of 58.0 percent and 69.2 percent respectively, with deepseek-v4-flash dropping to a minimum of 2.68 token/s. In contrast, minimax-m3 remains stable with a coefficient of variation of 14.6 percent. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage. Because the dataset lacks concurrent request volume metrics, throughput drops cannot be definitively attributed to capacity limits or low traffic.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 111.23 token/s, while nemotron-3-ultra is the weakest at 29.75 token/s. The most operationally significant volatility appears in deepseek-v4-pro and deepseek-v4-flash, which exhibit severe throughput drops to minimums of 2.76 token/s and 2.68 token/s respectively, alongside high coefficients of variation of 56.1% and 52.1%. In contrast, minimax-m3 delivers the most stable performance, averaging 43.98 token/s with a low coefficient of variation of 14.0%. The dataset contains 288 valid observations across six models, achieving 100.0% coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 113.14 token/s, while nemotron-3-ultra is the weakest at 35.2 token/s. The most operationally significant volatility comes from deepseek-v4-pro, which dropped from a high of 119.69 token/s to a low of 2.76 token/s, exhibiting a coefficient of variation of 56.1 percent and ending with a negative trend of 38.3 percent. In contrast, minimax-m3 remained stable with a coefficient of variation of 13.7 percent. The dataset includes 288 valid observations with 100.0 percent coverage, meaning there are no missing-data limitations to report for this period.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 107.88 token/s, while nemotron-3-ultra is the weakest at 42.4 token/s. The most operationally significant volatility appears in deepseek-v4-flash and nemotron-3-ultra, which exhibit extreme coefficient of variation values of 55.7% and 67.7% respectively. Deepseek-v4-flash throughput frequently collapses, dropping as low as 4.57 token/s, and nemotron-3-ultra similarly plunges to 3.26 token/s. In contrast, minimax-m3 delivers the most stable performance with a coefficient of variation of just 17.3% and an average of 45.31 token/s. The dataset contains 288 valid observations with 100.0% coverage, so there are no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 103.78 token/s, while nemotron-3-ultra is the weakest at 41.02 token/s. The most operationally significant volatility appears in deepseek-v4-flash and nemotron-3-ultra, which exhibit extreme coefficient of variation values of 67.2 percent and 76.4 percent respectively, with deepseek-v4-flash dropping to a minimum of 2.71 token/s. In contrast, minimax-m3 provides the most stable performance, maintaining an average of 44.83 token/s with a coefficient of variation of only 17.5 percent. The dataset contains complete coverage with no missing-data limitations across all 288 valid observations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the highest average throughput at 105.83 token/s, while nemotron-3-ultra was the weakest at 39.28 token/s. The most operationally significant volatility came from nemotron-3-ultra, which exhibited a 42.3 percent downward trend and a coefficient of variation of 80.8 percent, dropping to a minimum of 3.26 token/s. Similarly, deepseek-v4-flash showed extreme instability with a coefficient of variation of 64.5 percent and a low of 2.71 token/s. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage. However, the analysis is limited by the absence of concurrent request volume data, preventing correlation of throughput drops with traffic spikes.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model by average throughput at 106.55 token/s, while minimax-m3 is the weakest at 44.71 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which has a coefficient of variation of 79.2 percent and a maximum of 140.34 token/s but drops to a minimum of 3.26 token/s. Similarly, deepseek-v4-flash shows extreme instability, falling to 2.71 token/s despite a 127.11 token/s peak. In contrast, deepseek-v4-pro is the most stable, staying between 57.65 and 120.27 token/s. The dataset has 100.0 percent coverage with 288 valid points, so there are no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 107.77 token/s, while minimax-m3 was the weakest at 47.57 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which fluctuated heavily between a low of 3.97 token/s and a high of 140.34 token/s, yielding a coefficient of variation of 67.5 percent. deepseek-v4-flash also showed notable instability, dropping to 2.71 token/s late in the period. The dataset includes complete coverage with 288 valid points and no missing-data limitation.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 110.78 token/s, while minimax-m3 is the weakest at 50.71 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which has a coefficient of variation of 69.2 percent and throughput swings from a minimum of 6.32 token/s to a maximum of 140.34 token/s. Deepseek-v4-flash also shows instability, dropping to 3.85 token/s at 03:20. The dataset has 100.0 percent coverage with 288 valid points across all six models, so there are no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the highest average throughput at 106.97 token/s, while minimax-m3 was the weakest at 50.75 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which exhibited extreme swings between 4.4 and 158.35 token/s, yielding a coefficient of variation of 75.4 percent. Deepseek-v4-pro showed a notable upward trend of 41.9 percent, improving from early lows to sustained speeds above 117 token/s near the end of the period. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 111.63 token/s, while nemotron-3-ultra is the weakest at 44.04 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which has a coefficient of variation of 90.1 percent and swings between 4.14 token/s and 158.35 token/s. deepseek-v4-flash also shows notable instability, dropping to a low of 3.85 token/s despite an average of 92.48 token/s. The dataset includes 288 valid observations across six models with 100.0 percent coverage, so there are no missing-data limitations affecting this review.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 111.71 token/s, while nemotron-3-ultra is the weakest at 40.39 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits a 96.7% coefficient of variation with extreme swings between 4.14 and 158.35 token/s. Additionally, deepseek-v4-flash shows a severe downward trend, dropping 34.3% to a minimum of 3.85 token/s near the end of the period. The dataset contains 288 valid observations across all six models, achieving 100.0% coverage. Because the dataset is limited to this single rolling four-hour window, longer-term missing-data patterns outside this timeframe cannot be evaluated.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 114.55 token/s, closely followed by deepseek-v4-flash at 113.38 token/s. The weakest is nemotron-3-ultra at 45.7 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which has a coefficient of variation of 88.4 percent and drops to a minimum of 4.14 token/s. Deepseek-v4-pro shows a notable downward trend of 8.0 percent, ending at 69.31 token/s. The dataset includes 288 valid points with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Over the four-hour observation window, deepseek-v4-flash achieved the strongest average output-token throughput at 122.4 token/s, while nemotron-3-ultra was the weakest at 38.59 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which exhibited extreme instability with a coefficient of variation of 86.4 percent, a sharp downward trend of -36.9 percent, and a minimum drop to 4.14 token/s. In contrast, deepseek-v4-pro showed the most stable performance with a coefficient of variation of 18.6 percent. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash had the highest average throughput at 121.96 token/s, while nemotron-3-ultra was the weakest at 34.37 token/s. The most operationally significant volatility came from nemotron-3-ultra, which fluctuated wildly between 4.91 and 114.94 token/s with a 92.7% coefficient of variation, indicating severe instability. Conversely, deepseek-v4-pro offered the most stable performance, maintaining a 94.27 token/s average with only a 14.6% coefficient of variation. The dataset contains 288 valid observations across six models, achieving 100.0% coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash achieved the strongest average throughput at 119.26 token/s, while nemotron-3-ultra was the weakest at 27.57 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which fluctuated wildly between 3.51 and 114.94 token/s with a coefficient of variation of 113.6%, including a late-window spike. In contrast, minimax-m3 remained the most stable, averaging 52.29 token/s with a coefficient of variation of just 17.2%. The dataset includes 288 valid observations across six models, achieving 100% coverage. However, a missing-data limitation exists because the dataset lacks concurrent request counts, preventing per-request throughput normalization.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash is the strongest model with an average throughput of 119.83 token/s, while nemotron-3-ultra is the weakest at 13.88 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which fluctuated wildly between a minimum of 3.51 token/s and a maximum of 93.67 token/s, yielding a coefficient of variation of 138.3 percent. In contrast, minimax-m3 offered the most stable performance, averaging 52.9 token/s with a coefficient of variation of just 16.8 percent. The dataset includes complete observations with 100.0 percent coverage and no missing-data limitations across all 288 valid points.

4-hour window · 287 points · 99.7% coverage

Across the four-hour window, deepseek-v4-flash had the highest average throughput at 101.54 token/s, while nemotron-3-ultra was the weakest at 7.21 token/s. The most operationally significant volatility appeared in deepseek-v4-flash, which dropped to a minimum of 7.81 token/s despite reaching a maximum of 157.17 token/s. Similarly, gemma4:31b exhibited high instability with a coefficient of variation of 41.9 percent and a minimum of 20.26 token/s. In contrast, minimax-m3 provided the most stable performance, averaging 53.33 token/s with a low coefficient of variation of 14.0 percent. A missing-data limitation affects this review: deepseek-v4-pro has 47 valid samples instead of 48, dropping its coverage to 97.9 percent.

4-hour window · 287 points · 99.7% coverage

Across the four-hour window, deepseek-v4-flash achieved the highest average throughput at 93.81 token/s, while nemotron-3-ultra was the weakest at 14.38 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which experienced severe throughput drops to 7.81 token/s and 12.5 token/s before recovering to a peak of 157.17 token/s. Similarly, nemotron-3-ultra exhibited high instability, falling from 103.93 token/s to 0.35 token/s early in the period. Overall dataset coverage reached 99.7 percent across 287 valid points. However, the analysis is limited by a missing observation for deepseek-v4-pro, which recorded only 47 samples instead of the expected 48.

4-hour window · 287 points · 99.7% coverage

Across the four-hour window, deepseek-v4-flash had the strongest average throughput at 87.57 token/s, while nemotron-3-ultra was the weakest at 21.68 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which dropped sharply from a peak of 103.93 token/s down to single digits, ending at 6.22 token/s. Deepseek-v4-flash also exhibited extreme volatility, with throughput crashing to a minimum of 7.81 token/s before recovering. A missing-data limitation affects this review: deepseek-v4-pro has 47 samples instead of the expected 48, resulting in 97.9 percent coverage, whereas the other five models have complete data.

4-hour window · 287 points · 99.7% coverage

Across the four-hour window, glm-5.2 is the strongest model by average throughput at 82.87 token/s, while gemma4:31b is the weakest at 27.0 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which has a coefficient of variation of 92.1% and a downward trend of -31.5%, dropping from a maximum of 103.93 token/s to a latest reading of just 4.28 token/s. deepseek-v4-flash also shows severe instability, with throughput crashing to a minimum of 7.81 token/s. A missing-data limitation affects this review: deepseek-v4-pro has 47 valid samples instead of 48, resulting in 97.9% coverage. Overall dataset coverage is 99.7% across 287 valid points.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash achieved the strongest average throughput at 90.88 token/s, while gemma4:31b was the weakest at 22.44 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which exhibited extreme swings between 0.35 token/s and 130.51 token/s, yielding a coefficient of variation of 78.0 percent. Additionally, deepseek-v4-flash showed notable end-of-period instability, plummeting from 141.84 token/s at 17:40 to just 9.02 token/s by 18:00. The dataset includes 288 valid observations, achieving 100.0 percent coverage across all six models with no missing-data limitations.

4-hour window · 282 points · 97.9% coverage

Over the four-hour window, deepseek-v4-flash had the strongest average throughput at 94.69 token/s, while gemma4:31b was the weakest at 19.38 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which ranged from 6.7 to 130.51 token/s with a coefficient of variation of 70.5%, alongside deepseek-v4-flash dropping to 16.61 token/s at 13:25. This dataset contains a missing-data limitation: each of the six models recorded 47 valid samples out of 48 expected observations, yielding 97.9% coverage and leaving one five-minute interval unaccounted for across all models.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash had the strongest average output-token throughput at 92.52 token/s, while gemma4:31b was the weakest at 22.52 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which ranged from 6.7 to 130.51 token/s with a coefficient of variation of 76.4 percent, indicating highly unstable generation speeds. Similarly, deepseek-v4-flash experienced sharp drops to 16.61 token/s and 22.75 token/s despite its high average. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash had the strongest average output-token throughput at 91.11 token/s, while gemma4:31b was the weakest at 26.49 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which spiked to 130.51 token/s at 14:45 despite a low 31.74 token/s average, yielding a 75.8% coefficient of variation. Similarly, gemma4:31b exhibited a severe throughput drop around 13:35, falling to 5.27 token/s. The dataset contains 288 valid observations across six models, achieving 100.0% coverage with no missing-data limitations. This complete visibility confirms the observed fluctuations represent actual operational performance rather than monitoring gaps.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash shows the strongest average throughput at 90.54 token/s, while nemotron-3-ultra is the weakest at 28.45 token/s. The most operationally significant volatility appears in deepseek-v4-flash, which maintains a high average but exhibits steep drops, plunging to a minimum of 16.61 token/s at 13:25 before surging back to 133.89 token/s at 13:35. Similarly, gemma4:31b displays high instability with a 61.4% coefficient of variation and a downward trend of 23.9%, ending at 8.44 token/s. The dataset contains complete observations with no missing-data limitation, recording 288 valid points across the six models.

4-hour window · 288 points · 100.0% coverage

Over the four-hour window, deepseek-v4-flash had the strongest average output-token throughput at 84.14 token/s, while nemotron-3-ultra was weakest at 26.04 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which had a coefficient of variation of 67.2 percent and throughput swings between 5.89 and 69.14 token/s. deepseek-v4-flash also showed notable instability, dropping to 20.64 token/s before reaching 132.98 token/s. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitation.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 82.79 token/s, while nemotron-3-ultra was the weakest at 27.54 token/s. The most operationally significant volatility appeared in deepseek-v4-flash, which ranged from 13.15 to 132.98 token/s with a 43.8 percent coefficient of variation, indicating highly unstable output despite an 82.30 token/s average. Nemotron-3-ultra also showed severe swings, dropping to 5.89 token/s. The dataset contains 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitation. This complete visibility confirms that throughput fluctuations reflect actual system instability rather than monitoring gaps.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash had the strongest average output-token throughput at 86.16 token/s, while gemma4:31b was weakest at 23.88 token/s. The most operationally significant volatility came from nemotron-3-ultra, which dropped 45.8 percent over the period with a coefficient of variation of 76.2 percent, swinging between 104.14 token/s and 5.89 token/s. deepseek-v4-flash also showed high instability, plunging to 13.15 token/s before recovering. The dataset includes 288 valid observations across six models with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.