← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1207 retained summaries
4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 88.37 token/s, while minimax-m3 is the weakest at 39.32 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits extreme throughput swings, dropping to a minimum of 5.45 token/s before peaking at 152.72 token/s, yielding a coefficient of variation of 57.6 percent. This high variance contrasts with the steadier glm-5.2, which maintains a lower coefficient of variation of 23.2 percent. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 90.62 token/s, while minimax-m3 is the weakest at 41.15 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits a coefficient of variation of 57.1 percent and swings from a minimum of 5.45 token/s to a maximum of 152.72 token/s. Conversely, glm-5.2 is the most stable with a coefficient of variation of 26.0 percent. The dataset contains 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 89.85 token/s, while minimax-m3 is the weakest at 40.36 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which exhibits a 93.9 percent upward trend and a coefficient of variation of 89.4 percent, swinging from a 1.44 token/s minimum to a 152.72 token/s maximum. deepseek-v4-flash also shows notable instability with a 28.8 percent downward trend. The dataset contains 288 valid observations across all six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash had the strongest average output-token throughput at 93.14 token/s, while nemotron-3-ultra was weakest at 27.8 token/s. The most operationally significant volatility came from nemotron-3-ultra, which stayed near 2 token/s early but spiked to 112.31 token/s at 02:35, yielding a coefficient of variation of 121.2 percent. Deepseek-v4-flash also showed instability, dropping to 9.75 token/s at 03:45 despite its high average. The dataset contains 288 valid observations across six models with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Over the four-hour observation period, deepseek-v4-flash was the strongest model with an average throughput of 105.93 token/s, while nemotron-3-ultra was the weakest at 16.36 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which stayed below 13 token/s for over two hours before abruptly spiking to a maximum of 112.31 token/s, reflecting a coefficient of variation of 172.3 percent. Additionally, deepseek-v4-pro exhibited a sharp throughput drop, hitting a minimum of 4.28 token/s. The dataset includes 288 valid observations across all six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Over the four-hour observation period, deepseek-v4-flash was the strongest model, averaging 111.10 token/s, while nemotron-3-ultra was the weakest, averaging just 3.89 token/s. The most operationally significant volatility occurred in deepseek-v4-pro, which dropped sharply from 75.57 token/s at 01:50 to 4.28 token/s by 02:00. Similarly, gemma4:31b exhibited high volatility with a coefficient of variation of 31.6 percent, fluctuating between 42.23 and 130.77 token/s. The dataset includes 288 valid points across six models, achieving 100.0 percent coverage. Because the dataset lacks concurrent request volume metrics, it is impossible to determine if these throughput drops stem from capacity saturation or upstream demand changes.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash is the strongest model with an average throughput of 110.71 token/s, while nemotron-3-ultra is the weakest at 3.7 token/s. Operationally, minimax-m3 shows the most significant downward trend, dropping 17.5 percent to a latest throughput of 42.38 token/s, alongside high volatility with a coefficient of variation of 37.9 percent. In contrast, deepseek-v4-flash and deepseek-v4-pro trended upward by 10.5 percent and 10.0 percent, respectively. The dataset includes 288 valid observations across all six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Over the four-hour window, deepseek-v4-flash had the strongest average output-token throughput at 102.67 token/s, while nemotron-3-ultra was weakest at 8.22 token/s. The most operationally significant volatility was nemotron-3-ultra’s coefficient of variation of 129.5%, with throughput collapsing from a 47.28 token/s peak to sustained sub-5 token/s levels. Conversely, gemma4:31b showed a steep upward trend, climbing 68.9% to a late peak of 130.77 token/s. The dataset contains 288 valid points across six models, achieving 100.0% coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash had the strongest average output-token throughput at 98.04 token/s, while nemotron-3-ultra was weakest at 16.61 token/s. The most operationally significant volatility was nemotron-3-ultra’s severe degradation after 20:45 UTC, where throughput collapsed from roughly 44.71 token/s to 2.28 token/s, reflecting an 87.3 percent downward trend. Additionally, deepseek-v4-pro exhibited sharp instability between 20:20 and 20:35 UTC, dropping to a minimum of 8.78 token/s. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations, as all models maintained their expected 48 samples.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash had the strongest average throughput at 91.64 token/s, while nemotron-3-ultra was weakest at 22.46 token/s. The most operationally significant trend is nemotron-3-ultra's severe degradation, dropping 63.5 percent to a latest throughput of 4.34 token/s after 20:45. Additionally, deepseek-v4-pro exhibited high volatility, spiking to 121.52 token/s before crashing to 8.78 token/s around 20:35. The dataset records 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitations. This complete visibility confirms the observed throughput instabilities are genuine operational issues rather than reporting artifacts.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash is the strongest model with an average throughput of 94.0 token/s, while nemotron-3-ultra is the weakest at 28.15 token/s. The most operationally significant trend is minimax-m3 declining by 30.5 percent to a latest throughput of 31.33 token/s, alongside severe volatility for deepseek-v4-pro which dropped to a minimum of 8.78 token/s despite an average of 65.8 token/s. The dataset includes 288 valid observations with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash had the strongest average output-token throughput at 92.27 token/s, while nemotron-3-ultra was weakest at 29.9 token/s. The most operationally significant volatility occurred in deepseek-v4-pro, which dropped sharply to a minimum of 4.65 token/s at 17:00 despite an average of 63.84 token/s. minimax-m3 exhibited a notable downward trend, falling from 75.69 token/s at 16:10 to 14.84 token/s at 19:45. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash had the strongest average output-token throughput at 88.85 token/s, while nemotron-3-ultra was weakest at 25.89 token/s. The most operationally significant volatility occurred in deepseek-v4-pro, which dropped from 93.62 token/s at 15:55 to 4.65 token/s at 17:00, reflecting severe instability despite its 63.94 token/s average. Additionally, deepseek-v4-flash exhibited high volatility, swinging between 31.66 and 140.36 token/s. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations. This complete dataset allows for reliable operational monitoring of throughput performance across all tracked models.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash had the strongest average throughput at 84.03 token/s, while nemotron-3-ultra was weakest at 24.19 token/s. The most operationally significant volatility occurred in deepseek-v4-pro, which dropped sharply from 11.57 token/s at 16:55 to 4.65 token/s at 17:00 before recovering to 72.09 token/s at 17:05. Deepseek-v4-flash also exhibited high volatility, swinging between 20.26 and 136.84 token/s across the period. The dataset contains 288 valid observations with 100.0 percent coverage, matching the expected 48 samples per model, so there are no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour observation window, glm-5.2 delivered the strongest average output-token throughput at 78.58 token/s, while nemotron-3-ultra was the weakest at 23.45 token/s. The most operationally significant volatility occurred in deepseek-v4-pro, which dropped sharply from 81.83 token/s at 16:50 to 4.65 token/s by 17:00. Deepseek-v4-flash exhibited the highest single-interval peak, reaching 135.02 token/s at 15:45. Although the dataset reports 100.0 percent coverage across all six models, this analysis is limited by the absence of concurrent request volume data, preventing differentiation between model-side degradation and traffic-driven latency.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 78.7 token/s, while nemotron-3-ultra was the weakest at 23.49 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which ranged from a minimum of 17.57 token/s to a maximum of 135.02 token/s, yielding a high coefficient of variation of 36.2 percent. This erratic performance complicates capacity planning despite an upward trend of 10.8 percent. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 76.93 token/s, while nemotron-3-ultra is the weakest at 24.19 token/s. The most operationally significant trend is gemma4:31b's sharp decline of 46.6 percent, dropping from an early peak of 64.91 token/s to a latest reading of 21.15 token/s. Additionally, deepseek-v4-flash exhibits high volatility, with throughput swinging between 17.57 and 114.03 token/s. The dataset contains 288 valid observations across all six models, achieving 100.0 percent coverage. Because the data is limited to this single four-hour period, it lacks the historical context needed to determine if these throughput fluctuations represent normal diurnal patterns or isolated incidents.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash had the strongest average throughput at 77.56 token/s, while nemotron-3-ultra was weakest at 24.97 token/s. The most operationally significant volatility is deepseek-v4-flash's sharp late-window decline, dropping from 131.65 token/s at 10:30 to a minimum of 17.57 token/s by 14:00, a 24.9 percent downward trend. Similarly, glm-5.2 exhibited high volatility, spiking to 113.81 token/s at 13:55 after falling to 10.66 token/s at 11:35. The dataset includes 288 valid observations, achieving 100.0 percent coverage across all six models with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash achieved the highest average throughput at 78.04 token/s, while nemotron-3-ultra was the weakest at 26.04 token/s. The most operationally significant volatility appeared in minimax-m3, which fluctuated between 7.13 and 81.72 token/s despite a 44.44 token/s average, and nemotron-3-ultra, which dropped to 4.15 token/s. Deepseek-v4-pro showed a notable upward trend, rising 20.2 percent to a peak of 101.04 token/s. The dataset contains 288 valid observations across six models, yielding 100.0 percent coverage. Because the dataset lacks concurrent request counts, throughput drops cannot be attributed solely to model degradation.

4-hour window · 288 points · 100.0% coverage

Across the four-hour observation window, deepseek-v4-flash achieved the strongest average output-token throughput at 82.26 token/s, while nemotron-3-ultra was the weakest at 27.6 token/s. The most operationally significant trend is deepseek-v4-pro's sharp throughput increase, rising 115.8 percent from a low of 18.41 token/s early in the window to a peak of 101.04 token/s. High volatility also impacted nemotron-3-ultra, which dropped to a low of 4.15 token/s despite a 61.92 token/s maximum. The dataset includes 288 valid observations across six models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this performance review.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash had the strongest average throughput at 79.53 token/s, while nemotron-3-ultra was weakest at 27.85 token/s. The most operationally significant trend was deepseek-v4-pro's recovery from a low of 18.41 token/s at 09:00 to 101.04 token/s at 10:40, reflecting an 18.5 percent upward trend. High volatility was also evident, as glm-5.2 exhibited a sharp drop from 117.29 token/s at 09:05 to 18.02 token/s at 09:35. The dataset contains complete coverage with 288 valid points and no missing-data limitation across the six models.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 72.64 token/s, while nemotron-3-ultra was the weakest at 29.32 token/s. The most operationally significant trend is the severe throughput degradation in deepseek-v4-pro, which declined by 45.6 percent to a late-window average near 25 token/s, contrasting with deepseek-v4-flash, which trended upward by 13.1 percent. All six models exhibited high volatility, with coefficients of variation ranging from 35.0 to 48.2 percent and frequent drops below 20 token/s. The dataset contains complete coverage with 288 valid observations, so there are no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the highest average output-token throughput at 80.13 token/s, while nemotron-3-ultra was the weakest at 30.46 token/s. The most operationally significant volatility appeared in deepseek-v4-flash, which fluctuated between a low of 15.11 token/s and a high of 132.92 token/s, alongside a 38.8 coefficient of variation percentage. Additionally, deepseek-v4-pro exhibited a notable downward trend, dropping 25.1 percent to a latest throughput of 18.41 token/s. The dataset includes 288 valid observations, achieving 100.0 percent coverage across all six models with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 85.39 token/s, while nemotron-3-ultra was the weakest at 28.32 token/s. The most operationally significant trend was the sharp throughput decline in deepseek-v4-flash, which fell by 21.5 percent to a latest reading of 93.79 token/s despite maintaining a high average of 76.33 token/s. Additionally, nemotron-3-ultra exhibited high volatility, dropping to a minimum of 3.1 token/s. The dataset contains 288 valid observations across all six models, achieving 100.0 percent coverage with no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 85.92 token/s, while nemotron-3-ultra is the weakest at 29.86 token/s. The most operationally significant volatility appears in deepseek-v4-flash, which maintains a high average of 80.76 token/s but exhibits severe swings, dropping to a minimum of 15.11 token/s and showing a 17.0 percent downward trend. Similarly, nemotron-3-ultra experiences extreme instability, plunging to 3.1 token/s. The dataset contains 288 valid observations with 100.0 percent coverage, meaning there are no missing-data limitations to constrain this operational review.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash had the strongest average output-token throughput at 87.95 token/s, while nemotron-3-ultra was weakest at 32.02 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which dropped to a 3.1 token/s minimum despite a 58.94 token/s maximum, yielding a 49.9% coefficient of variation. Additionally, gemma4:31b exhibited a notable downward trend, falling 29.8% to a latest throughput of 25.94 token/s. The dataset contains 288 valid observations across six models, achieving 100.0% coverage. However, a missing-data limitation exists: the dataset lacks concurrent request concurrency counts, preventing correlation of throughput drops with specific load conditions.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash had the strongest average throughput at 88.7 token/s, while nemotron-3-ultra was weakest at 36.63 token/s. The most operationally significant trend is nemotron-3-ultra's 23.4 percent decline, dropping from 62.19 token/s at 00:05 to 17.06 token/s at 04:00, alongside severe volatility that included a plunge to 5.01 token/s at 02:20. deepseek-v4-pro also degraded significantly, trending down 14.8 percent to a late low of 11.59 token/s at 03:10. The dataset records 288 valid points across all models with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 87.1 token/s, while nemotron-3-ultra was the weakest at 38.49 token/s. The most operationally significant volatility came from gemma4:31b, which fluctuated heavily between a minimum of 25.71 token/s and a maximum of 123.93 token/s, yielding a coefficient of variation of 43.0 percent. Deepseek-v4-pro showed a notable upward trend of 21.9 percent over the period. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage. Because the dataset lacks request concurrency and prompt length metadata, throughput drops cannot be attributed to load spikes or input size.

4-hour window · 288 points · 100.0% coverage

Over the four-hour observation window, deepseek-v4-flash delivered the strongest average output-token throughput at 84.25 token/s, while nemotron-3-ultra was the weakest at 43.03 token/s. The most operationally significant volatility occurred in gemma4:31b, which exhibited severe throughput instability, swinging from a maximum of 135.28 token/s down to a minimum of 25.71 token/s with a high coefficient of variation of 43.1 percent. This erratic performance contrasts with minimax-m3, which maintained much steadier output around its 54.09 token/s average. All six models achieved complete data coverage with 48 out of 48 expected samples collected, meaning there are no missing-data limitations affecting this specific rolling window analysis.

4-hour window · 288 points · 100.0% coverage

Over the four-hour observation period, deepseek-v4-flash achieved the strongest average output-token throughput at 82.82 token/s, while nemotron-3-ultra was the weakest at 38.82 token/s. The most operationally significant volatility came from nemotron-3-ultra, which exhibited extreme instability with a coefficient of variation of 52.6% and throughput dropping as low as 2.72 token/s. Conversely, minimax-m3 provided the most stable performance, maintaining a steady 56.39 token/s average with a low coefficient of variation of 14.7%. The dataset includes 288 valid observations across six models, achieving 100.0% coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-pro had the strongest average output-token throughput at 82.1 token/s, while nemotron-3-ultra was weakest at 35.61 token/s. The most operationally significant volatility came from nemotron-3-ultra, which dropped to a minimum of 2.72 token/s and exhibited a coefficient of variation of 57.7 percent, indicating highly unstable generation speeds. Conversely, minimax-m3 provided the most stable performance, varying only by 8.32 token/s standard deviation. All six models maintained complete dataset coverage with 48 out of 48 expected samples collected, meaning there are no missing-data limitations affecting this specific analysis period.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-pro had the strongest average output-token throughput at 81.87 token/s, while nemotron-3-ultra was weakest at 34.49 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which dropped sharply to a minimum of 2.72 token/s around 21:25 before recovering. glm-5.2 also exhibited extreme instability, fluctuating between 5.6 and 124.65 token/s. In contrast, minimax-m3 provided the most stable performance, maintaining a steady 55.39 token/s average with a low coefficient of variation. The dataset includes complete observations with no missing-data limitation, capturing all 288 expected points across the six models.

4-hour window · 288 points · 100.0% coverage

Across the four-hour observation window, deepseek-v4-flash showed the strongest average output-token throughput at 76.97 token/s, while nemotron-3-ultra was the weakest at 29.52 token/s. The most operationally significant volatility occurred in glm-5.2, which exhibited extreme swings despite a 69.03 token/s average, plummeting to a minimum of 5.6 token/s before surging to 113.36 token/s. Conversely, minimax-m3 delivered the most stable performance, maintaining a 53.63 token/s average with a low coefficient of variation of 13.7 percent. The dataset contains 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the highest average output-token throughput at 68.84 token/s, while nemotron-3-ultra was the weakest at 31.70 token/s. The most operationally significant volatility appears in deepseek-v4-flash, which swung from a low of 3.53 token/s to a high of 137.04 token/s, reflecting a coefficient of variation of 49.0 percent. Similarly, glm-5.2 experienced severe drops, falling to 5.60 token/s at 19:40. In contrast, minimax-m3 maintained the most stable performance, varying only between 35.12 and 67.47 token/s. The dataset includes complete observations with no missing-data limitations, covering all 288 expected samples.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 72.3 token/s, while gemma4:31b was the weakest at 28.59 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which exhibited extreme instability with a coefficient of variation of 63.3 percent, dropping to a minimum of 3.53 token/s before surging to a peak of 137.04 token/s. This high variance contrasts with minimax-m3, which maintained the most stable performance at 52.17 token/s on average. The dataset contains 288 valid observations across all six models, achieving 100.0 percent coverage with no missing data limitations affecting the throughput analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 75.43 token/s, while nemotron-3-ultra was the weakest at 28.25 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which exhibited extreme instability with a coefficient of variation of 76.9 percent, dropping to a minimum of 3.53 token/s before surging to a maximum of 105.83 token/s. This model also showed a drastic upward trend of 107.2 percent over the period. The dataset contains no missing-data limitations, as all six models maintained complete coverage with exactly 48 valid samples each, matching the expected count perfectly.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 78.64 token/s, while deepseek-v4-flash was the weakest at 25.56 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which exhibited extreme instability with a coefficient of variation of 81.2 percent, dropping to a minimum of 3.53 token/s before spiking to 101.76 token/s. Similarly, deepseek-v4-pro showed high volatility with a standard deviation of 25.45 token/s. The dataset contains 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitations, providing a complete view of throughput performance for all monitored models.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the highest average output-token throughput at 72.27 token/s, while deepseek-v4-flash was the weakest at 23.99 token/s. The most operationally significant volatility occurred in deepseek-v4-pro, which dropped sharply from 98.35 token/s at 13:35 to 7.16 token/s at 14:05, reflecting severe throughput instability. Additionally, glm-5.2 exhibited a strong upward trend, increasing by 40.2 percent over the period and peaking at 117.3 token/s near 16:45. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 62.55 token/s, while nemotron-3-ultra was the weakest at 31.48 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which dropped sharply from early highs near 116.36 token/s to a low of 10.42 token/s, reflecting a trend decline of 68.4 percent. Similarly, gemma4:31b exhibited high variability with a coefficient of variation of 48.0 percent and a 34.1 percent downward trend. In contrast, minimax-m3 remained stable with an average of 52.61 token/s and a coefficient of variation of just 14.4 percent. The dataset includes complete coverage with no missing-data limitations across all 288 valid observations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 64.42 token/s, while nemotron-3-ultra was the weakest at 31.03 token/s. The most operationally significant volatility appears in deepseek-v4-flash, which dropped from a maximum of 135.87 token/s to a low of 15.23 token/s with a coefficient of variation of 52.8 percent, alongside a negative trend of 37.9 percent ending at 16.87 token/s. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitation.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 75.13 token/s, while nemotron-3-ultra is the weakest at 29.14 token/s. The most operationally significant volatility appears in glm-5.2 and deepseek-v4-flash, which exhibit sharp throughput drops to 12.83 token/s and 17.9 token/s respectively, alongside high coefficients of variation around 37 to 39 percent. In contrast, minimax-m3 maintains the most stable performance, averaging 54.39 token/s with a low coefficient of variation of 13.9 percent. The dataset contains 288 valid observations across six models, achieving 100 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 81.32 token/s, while nemotron-3-ultra is the weakest at 30.38 token/s. The most operationally significant volatility appears in deepseek-v4-flash, which oscillates between 17.9 and 135.87 token/s despite a steady average of 74.15 token/s. Similarly, glm-5.2 and deepseek-v4-pro both exhibit severe mid-window drops, with throughput plummeting to 12.83 and 11.23 token/s respectively. In contrast, minimax-m3 remains stable with a low coefficient of variation of 14.9 percent. The dataset contains no missing-data limitation, achieving 100.0 percent coverage across all 288 valid observations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 83.99 token/s, while gemma4:31b was the weakest at 31.01 token/s. The most operationally significant volatility appeared in deepseek-v4-flash, which ranged from 15.79 to 135.87 token/s with a coefficient of variation of 40.1 percent, indicating highly unstable output capacity. In contrast, minimax-m3 maintained the steadiest performance, fluctuating only between 38.6 and 71.53 token/s. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations. All expected samples were recorded successfully throughout the monitoring period.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average output-token throughput at 82.3 token/s, while nemotron-3-ultra was the weakest at 28.6 token/s. The most operationally significant volatility appears in deepseek-v4-flash, which ranged from 15.79 to 127.74 token/s with a coefficient of variation of 36.7 percent, indicating highly unstable generation speeds despite a high 75.94 token/s average. Similarly, gemma4:31b experienced severe drops, hitting a low of 8.24 token/s. In contrast, minimax-m3 maintained the most stable performance, fluctuating only between 38.6 and 70.2 token/s. The dataset includes 288 valid observations with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash had the strongest average output-token throughput at 80.84 token/s, while nemotron-3-ultra was weakest at 26.43 token/s. The most operationally significant volatility came from nemotron-3-ultra, which dropped from 42.86 token/s at 05:15 to 2.53 token/s at 06:15, reflecting a 70.4 coefficient of variation. Additionally, gemma4:31b exhibited a sharp downward trend, falling from a peak of 103.54 token/s at 05:55 to a low of 8.24 token/s at 08:05. The dataset contains 288 valid observations across all six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 282 points · 97.9% coverage

Across the four-hour window, glm-5.2 delivered the highest average throughput at 80.98 token/s, while nemotron-3-ultra was the weakest at 26.96 token/s. The most operationally significant volatility occurred in gemma4:31b, which experienced a severe throughput decline, dropping from a peak of 126.71 token/s down to 9.0 token/s before partially recovering, resulting in a 47.1 percent downward trend. Nemotron-3-ultra also showed extreme instability, plunging to 2.53 token/s for over twenty minutes. Analysis is limited by missing data, as each model recorded 47 samples instead of the expected 48, yielding 97.9 percent coverage.

4-hour window · 288 points · 100.0% coverage

Over the four-hour period, glm-5.2 is the strongest model with an average throughput of 83.98 token/s, closely followed by deepseek-v4-flash at 83.59 token/s. The weakest model is nemotron-3-ultra at 25.73 token/s. Operationally, nemotron-3-ultra shows the most significant volatility and degradation, dropping steadily from 38.14 token/s at 05:05 to a minimum of 2.53 token/s at 06:15, reflecting a 47.2 percent negative trend. In contrast, deepseek-v4-pro improved by 20.5 percent. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 delivered the strongest average throughput at 87.04 token/s, while nemotron-3-ultra was the weakest at 28.94 token/s. The most operationally significant volatility appeared in gemma4:31b, which fluctuated heavily between a minimum of 15.1 token/s and a maximum of 126.71 token/s, yielding a coefficient of variation of 47.9 percent. Similarly, nemotron-3-ultra exhibited extreme instability, dropping to just 2.9 token/s at 06:00 UTC. The dataset contains no missing-data limitation, as all six models maintained a 100.0 percent coverage rate across the expected 48 samples per model.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model with an average throughput of 91.94 token/s, while nemotron-3-ultra is the weakest at 34.27 token/s. The most operationally significant volatility appears in gemma4:31b, which has a coefficient of variation of 51.9 percent and throughput swings from a minimum of 15.1 token/s to a maximum of 132.98 token/s. This high volatility occurs despite a complete dataset with 100.0 percent coverage and 288 valid points across all six models. However, this four-hour dataset lacks prior baseline data, limiting the ability to determine if these throughput fluctuations represent normal operational behavior or emerging infrastructure degradation.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, glm-5.2 is the strongest model by average throughput at 90.88 token/s, closely followed by deepseek-v4-flash at 91.45 token/s. The weakest is nemotron-3-ultra at 33.9 token/s. Operationally, gemma4:31b shows the most significant volatility, with a high coefficient of variation of 49.8 percent and throughput dropping from a maximum of 132.98 token/s to a minimum of 15.1 token/s. It also exhibits a steep downward trend of -31.9 percent over the period. The dataset has 100.0 percent coverage with 288 valid points, so there are no missing-data limitations affecting this analysis.