← Performance dashboard

Hourly performance insights

Summaries of rolling four-hour performance data

1203 retained summaries
4-hour window · 336 points · 100.0% coverage

Across the four-hour window, gemma4:31b delivered the strongest average throughput at 122.15 token/s, while nemotron-3-ultra was the weakest at 21.86 token/s. The most operationally significant volatility comes from glm-5.3-flash, which dropped sharply from 154.48 token/s at 21:20 to 52.44 token/s at 21:30, driving a 37.3 percent downward trend. Nemotron-3-ultra also showed extreme instability, fluctuating between 0.99 and 85.99 token/s with a coefficient of variation of 115.1 percent. The dataset contains 336 valid observations across seven models with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.

4-hour window · 336 points · 100.0% coverage

Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 123.12 token/s, while nemotron-3-ultra is the weakest at 21.18 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits a coefficient of variation of 102.7 percent and drops to a minimum of 0.99 token/s. In contrast, deepseek-v4-pro is the most stable with a coefficient of variation of 14.6 percent. The dataset contains 336 valid observations across seven models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 336 points · 100.0% coverage

Across the four-hour window, glm-5.3-flash is the strongest model with an average throughput of 126.23 token/s, while nemotron-3-ultra is the weakest at 20.16 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits a coefficient of variation of 108.0 percent and drops to a minimum of 0.99 token/s at 20:05. Additionally, deepseek-v4-flash shows notable volatility with a minimum of 13.12 token/s despite an average of 99.88 token/s. The dataset contains 336 valid points across seven models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 336 points · 100.0% coverage

Across the four-hour window, glm-5.3-flash had the strongest average output-token throughput at 119.06 token/s, while nemotron-3-ultra was weakest at 14.14 token/s. The most operationally significant volatility appeared in deepseek-v4-flash, which dropped to a minimum of 13.12 token/s despite an average of 94.22 token/s, yielding a coefficient of variation of 30.6 percent. Similarly, gemma4:31b experienced a severe single-interval drop to 23.32 token/s against its 117.8 token/s average. The dataset contains 336 valid points across seven models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 336 points · 100.0% coverage

Across the four-hour observation window, gemma4:31b delivered the strongest average output-token throughput at 115.76 token/s, while nemotron-3-ultra was the weakest at 10.19 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which exhibited extreme instability with a coefficient of variation of 96.9 percent and throughput swings ranging from 1.63 to 46.49 token/s. Conversely, deepseek-v4-pro maintained the most stable performance, averaging 107.18 token/s with a low coefficient of variation of 11.2 percent. The dataset contains 336 valid observations across seven models, achieving complete coverage with no missing-data limitations.

4-hour window · 336 points · 100.0% coverage

Across the four-hour window, gemma4:31b delivered the strongest average output-token throughput at 109.36 token/s, while nemotron-3-ultra was the weakest at 7.26 token/s. The most operationally significant volatility came from glm-5.2, which exhibited a 34.5% upward trend alongside a 31.3% coefficient of variation, swinging from a low of 23.18 token/s to a high of 126.49 token/s. In contrast, deepseek-v4-pro maintained the most stable performance, varying only 11.9% around its 105.42 token/s average. The dataset contains 336 valid observations across seven models, achieving 100.0% coverage with no missing-data limitations.

4-hour window · 336 points · 100.0% coverage

Over the four-hour window, deepseek-v4-pro had the strongest average throughput at 105.45 token/s, while nemotron-3-ultra was weakest at 4.51 token/s. The most operationally significant volatility appeared in glm-5.2, which swung from a low of 23.18 token/s to a high of 126.49 token/s with a 33.0% coefficient of variation, and deepseek-v4-flash, which dropped sharply to 36.72 token/s at 15:45. In contrast, deepseek-v4-pro remained stable with a 12.0% coefficient of variation. The dataset includes 336 valid observations across seven models, achieving 100.0% coverage with no missing-data limitations.

4-hour window · 336 points · 100.0% coverage

Across the four-hour observation window, glm-5.3-flash achieved the strongest average output-token throughput at 105.25 token/s, while nemotron-3-ultra was the weakest at 3.71 token/s. The most operationally significant volatility occurred in glm-5.2, which exhibited a high coefficient of variation of 33.4% and sharp drops to a minimum of 23.18 token/s. In contrast, deepseek-v4-pro maintained the most stable performance with a coefficient of variation of just 12.5% and an average of 102.33 token/s. The dataset contains 336 valid points across seven models, achieving 100.0% coverage. However, analysis is limited by the absence of concurrent request volume data, preventing correlation of throughput drops with load spikes.

4-hour window · 336 points · 100.0% coverage

Across the four-hour observation window, glm-5.3-flash was the strongest model by average throughput at 110.91 token/s, while nemotron-3-ultra was the weakest at 4.3 token/s. The most operationally significant trend is the sharp throughput decline in glm-5.2, which dropped 23.5 percent overall and hit a low of 23.18 token/s near 14:15. In contrast, deepseek-v4-pro showed the most stable performance, maintaining a low coefficient of variation of 14.0 percent and an average of 100.01 token/s. The dataset includes 336 valid observations across seven models with 100.0 percent coverage, so there are no missing-data limitations affecting this specific interval.

4-hour window · 336 points · 100.0% coverage

Across the four-hour observation window, glm-5.3-flash achieved the strongest average output-token throughput at 118.49 token/s, while nemotron-3-ultra was the weakest at 4.65 token/s. The most operationally significant trend was a broad throughput decline near the period end, with glm-5.2 dropping 22.0 percent and nemotron-3-ultra dropping 30.5 percent. Additionally, glm-5.2 exhibited high volatility, ranging from a minimum of 27.58 token/s to a maximum of 135.96 token/s. The dataset includes 336 valid observations across seven models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this specific analysis.

4-hour window · 336 points · 100.0% coverage

Across the four-hour observation window, glm-5.3-flash was the strongest model with an average throughput of 124.90 token/s, while nemotron-3-ultra was the weakest at 4.61 token/s. The most operationally significant volatility occurred in glm-5.2, which experienced sharp throughput drops, falling from 116.55 token/s at 10:25 to 35.80 token/s at 10:30, and later from 97.32 token/s at 12:15 to 31.22 token/s at 12:20. Deepseek-v4-pro exhibited the most stable performance with a coefficient of variation of 13.7 percent. The dataset includes complete coverage with 336 valid points across all models, so there are no missing-data limitations affecting this analysis.

4-hour window · 336 points · 100.0% coverage

Across the four-hour observation window, glm-5.3-flash was the strongest model with an average throughput of 120.05 token/s, while nemotron-3-ultra was the weakest at 4.34 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which exhibited a 48.0% coefficient of variation and a maximum throughput of only 10.95 token/s despite a 72.2% upward trend percentage. In contrast, deepseek-v4-pro offered the most stable performance, maintaining a 99.07 token/s average with a 14.9% coefficient of variation. The dataset contains 336 valid observations across seven models with 100.0% coverage, meaning there are no missing-data limitations affecting this specific analysis period.

4-hour window · 325 points · 96.7% coverage

Across the four-hour window, glm-5.3-flash had the strongest average throughput at 120.97 token/s, while nemotron-3-ultra was the weakest at 3.93 token/s. Operationally, gemma4:31b showed high volatility, swinging from a low of 48.24 token/s to a maximum of 150.6 token/s with a coefficient of variation of 25.8 percent. Additionally, glm-5.3-flash exhibited a sharp upward trend, increasing 22.0 percent over the period. A missing-data limitation affects glm-5.3-flash, which has only 37 valid samples out of 48 expected, resulting in 77.1 percent coverage compared to the complete 48 samples for all other models.

4-hour window · 313 points · 93.2% coverage

Over the four-hour window, gemma4:31b delivered the strongest average throughput at 115.06 token/s, while nemotron-3-ultra was the weakest at 3.9 token/s. The most operationally significant trend is nemotron-3-ultra declining 31.0 percent to a latest throughput of 3.24 token/s, indicating a severe degradation. Additionally, gemma4:31b exhibited high volatility, dropping from a maximum of 165.59 token/s to a minimum of 48.24 token/s. This analysis is constrained by a missing-data limitation for glm-5.3-flash, which only collected 25 of 48 expected samples, yielding 52.1 percent coverage and limiting visibility into its 112.04 token/s average.

4-hour window · 301 points · 89.6% coverage

Over the four-hour window, gemma4:31b is the strongest model with an average throughput of 115.53 token/s, while nemotron-3-ultra is the weakest at 8.18 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits a coefficient of variation of 128.5 percent and a downward trend of 71.1 percent, dropping from a maximum of 58.32 token/s to a latest reading of 2.48 token/s. In contrast, minimax-m3 shows high stability with a coefficient of variation of 17.9 percent and an average of 42.52 token/s. A key missing-data limitation affects glm-5.3-flash, which has only 13 valid samples out of 48 expected, restricting visibility into its true performance.

4-hour window · 282 points · 83.9% coverage

Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 111.27 token/s, while nemotron-3-ultra is the weakest at 16.32 token/s. The most operationally significant trend is nemotron-3-ultra's severe degradation, dropping 83.5 percent from an early peak of 79.23 token/s down to 2.59 token/s. Additionally, deepseek-v4-pro exhibits high volatility with a coefficient of variation of 32.5 percent, including steep drops to 10.88 token/s. This analysis is limited by missing data: glm-5.3-flash has zero samples, and overall coverage is 83.9 percent with 282 valid points out of 336 expected.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash achieved the highest average throughput at 109.47 token/s, while nemotron-3-ultra was the weakest at 22.71 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which exhibited an 85.4% coefficient of variation and a severe downward trend of 61.2%, dropping from a peak of 79.23 token/s to just 5.12 token/s by 07:00. Deepseek-v4-pro also showed notable instability, with throughput crashing to 10.88 token/s at 04:55. The dataset includes 288 valid observations across six models, achieving 100.0% coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash had the highest average throughput at 107.29 token/s, while nemotron-3-ultra was the weakest at 27.72 token/s. The most operationally significant volatility appeared in deepseek-v4-pro, which dropped to a minimum of 10.88 token/s despite an average of 88.64 token/s, and nemotron-3-ultra, which fell to 4.7 token/s. In contrast, minimax-m3 was the most stable, averaging 43.23 token/s with a low coefficient of variation of 20.0 percent. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, gemma4:31b delivered the highest average output-token throughput at 110.36 token/s, while nemotron-3-ultra was the weakest at 24.34 token/s. The most operationally significant volatility came from nemotron-3-ultra, which exhibited extreme instability with a coefficient of variation of 78.8 percent and a minimum drop to 0.76 token/s. Conversely, deepseek-v4-pro showed a steep positive trend of 60.3 percent, climbing from early lows to a 130.65 token/s peak. The dataset contains 288 valid observations across all six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, gemma4:31b had the strongest average throughput at 116.55 token/s, while nemotron-3-ultra was weakest at 23.19 token/s. The most operationally significant volatility was a severe glm-5.2 throughput collapse between 01:10 and 01:30 UTC, where output dropped from 13.19 token/s to a minimum of 3.91 token/s before recovering. Deepseek-v4-pro also exhibited extreme instability, with a coefficient of variation of 49.4 percent and frequent swings between 21.99 token/s and 130.65 token/s. The dataset includes 288 valid observations across all six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, gemma4:31b delivered the strongest average throughput at 119.27 token/s, while nemotron-3-ultra was the weakest at 17.82 token/s. The most operationally significant volatility occurred in glm-5.2, which experienced a severe throughput collapse around 01:15 UTC, dropping from 120.23 token/s at 00:10 to a minimum of 3.91 token/s before partially recovering. Deepseek-v4-pro also exhibited extreme instability, fluctuating between 121.97 token/s and 21.99 token/s throughout the period. The dataset contains 288 valid observations across all six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, gemma4:31b delivered the strongest average output-token throughput at 123.15 token/s, while nemotron-3-ultra was the weakest at 11.97 token/s. The most operationally significant volatility occurred in glm-5.2, which experienced a severe throughput collapse between 01:10 and 01:30 UTC, dropping from 13.19 token/s to a low of 3.91 token/s. Deepseek-v4-pro also showed high instability, trending downward by 37.7 percent to a late reading of 25.34 token/s. In contrast, deepseek-v4-flash maintained the most stable performance with a coefficient of variation of 14.5 percent. The dataset includes 288 valid observations with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Over the four-hour window, deepseek-v4-flash had the strongest average throughput at 120.0 token/s, while nemotron-3-ultra was weakest at 15.99 token/s. The most operationally significant volatility came from nemotron-3-ultra, which had a coefficient of variation of 121.4 percent and dropped to 1.02 token/s. Deepseek-v4-pro also showed instability, falling to 23.19 token/s despite an 84.73 token/s average. Deepseek-v4-flash and glm-5.2 trended upward by 18.9 percent and 16.8 percent respectively. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash is the strongest model with an average throughput of 116.92 token/s, while nemotron-3-ultra is the weakest at 13.0 token/s. The most operationally significant trend is nemotron-3-ultra's severe degradation, dropping from a peak of 66.02 token/s to near-zero rates around 1.0 token/s before a late recovery, yielding a negative trend of 73.1 percent. Additionally, deepseek-v4-pro exhibits high volatility with a 33.4 percent coefficient of variation and a minimum of 5.11 token/s. The dataset contains 288 valid observations with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, gemma4:31b is the strongest model by average throughput at 114.8 token/s, while nemotron-3-ultra is the weakest at 14.68 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits a coefficient of variation of 116.9 percent and a downward trend of -28.1 percent, plummeting from a peak of 66.02 token/s to just 1.89 token/s by the end. Additionally, deepseek-v4-pro experienced a severe throughput drop to a minimum of 5.11 token/s around 20:35. The dataset contains 288 valid observations across all six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash had the strongest average output-token throughput at 108.46 token/s, while nemotron-3-ultra was weakest at 15.84 token/s. The most operationally significant volatility was nemotron-3-ultra’s extreme instability, with a coefficient of variation of 103.9 percent and throughput collapsing from a peak of 66.02 token/s down to 1.73 token/s. Additionally, deepseek-v4-pro experienced a severe throughput drop to 5.11 token/s at 20:35. The dataset contains 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 111.52 token/s, while nemotron-3-ultra is the weakest at 11.45 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits extreme instability with a coefficient of variation of 114.1 percent and erratic spikes up to 66.02 token/s. Additionally, deepseek-v4-pro experienced a severe throughput drop to 5.11 token/s at 20:35. The dataset includes 288 valid observations across six models, achieving complete coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 112.89 token/s, while nemotron-3-ultra is the weakest at 11.07 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which has a coefficient of variation of 103.5 percent and drops to a minimum of 2.09 token/s. In contrast, deepseek-v4-pro shows the most stability with a coefficient of variation of 17.2 percent and an average of 86.3 token/s. The dataset includes 288 valid points across six models with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 104.22 token/s, while nemotron-3-ultra is the weakest at 8.11 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits a 109.8% coefficient of variation and a -44.0% trend, dropping to a 0.4 token/s minimum despite a 51.96 token/s peak. In contrast, minimax-m3 is the most stable, averaging 43.76 token/s with an 18.9% coefficient of variation. The dataset contains 288 valid points across six models, achieving 100.0% coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Over the four-hour window, gemma4:31b had the strongest average throughput at 110.15 token/s, while nemotron-3-ultra was the weakest at 7.19 token/s. The most operationally significant volatility came from nemotron-3-ultra, which fluctuated wildly between 0.4 and 51.96 token/s with a coefficient of variation of 129.6 percent. In contrast, minimax-m3 was the most stable, averaging 44.62 token/s with a coefficient of variation of 20.9 percent. Deepseek-v4-flash and glm-5.2 showed moderate instability, frequently dropping below 40 token/s. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 287 points · 99.7% coverage

Over the four-hour window, gemma4:31b had the strongest average throughput at 106.84 token/s, while nemotron-3-ultra was the weakest at 12.89 token/s. The most operationally significant volatility was nemotron-3-ultra collapsing from 79.73 token/s at 13:05 to 0.40 token/s at 15:45, driving its coefficient of variation to 130.7 percent. Conversely, glm-5.2 showed a strong upward trend, increasing 22.7 percent to finish at 63.97 token/s. The dataset contains 287 valid points out of 288 expected across all models, yielding 99.7 percent coverage. This analysis is limited by one missing minimax-m3 observation at 13:15, leaving that model with 47 samples.

4-hour window · 287 points · 99.7% coverage

Over the four-hour window, gemma4:31b had the strongest average throughput at 101.92 token/s, while nemotron-3-ultra was weakest at 15.98 token/s. The most operationally significant trend was nemotron-3-ultra's severe degradation, dropping from a peak of 79.73 token/s at 13:05 to near-zero levels around 0.4 token/s by 15:45, reflecting an 87.9 percent downward trend. Additionally, deepseek-v4-flash exhibited high volatility, spiking to 131.27 token/s at 15:55 after dipping to 15.88 token/s at 14:45. A missing-data limitation affects this review: minimax-m3 has 47 samples instead of 48, missing the 13:15 UTC observation, yielding 97.9 percent coverage compared to the complete data for the other models.

4-hour window · 281 points · 97.6% coverage

Over the four-hour window, gemma4:31b is the strongest model with an average throughput of 106.9 token/s, while nemotron-3-ultra is the weakest at 20.45 token/s. The most operationally significant trend is nemotron-3-ultra's severe degradation, dropping from 79.73 token/s at 13:05 to near zero around 14:20, reflecting a 36.1 percent downward trend. Additionally, deepseek-v4-flash exhibits high volatility with a 33.1 percent coefficient of variation, including sharp drops to 13.48 token/s and 15.88 token/s. A missing-data limitation affects the dataset: minimax-m3 has 46 samples instead of 48, entirely missing observations at 13:15 and 13:20, resulting in 95.8 percent coverage.

4-hour window · 287 points · 99.7% coverage

Across the four-hour window, gemma4:31b had the strongest average output-token throughput at 101.17 token/s, while nemotron-3-ultra was the weakest at 28.42 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which exhibited extreme instability with a coefficient of variation of 78.6 percent and a sharp drop to a low of 2.45 token/s. Additionally, deepseek-v4-pro showed a notable downward trend, declining 22.6 percent over the period. A missing-data limitation affects this review: minimax-m3 recorded only 47 valid samples out of 48 expected, resulting in 97.9 percent coverage and leaving a gap in its continuous monitoring.

4-hour window · 288 points · 100.0% coverage

Across the four-hour observation period, gemma4:31b is the strongest model with an average throughput of 100.98 token/s, while nemotron-3-ultra is the weakest at 31.74 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits a coefficient of variation of 75.7 percent and a downward trend of 36.1 percent, dropping from an early peak of 110.59 token/s to a low of 4.98 token/s. This high volatility indicates severe throughput instability. The dataset includes 288 valid points across six models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 100.95 token/s, while nemotron-3-ultra is the weakest at 34.95 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits a 74.3% coefficient of variation and a downward trend of 31.9%, dropping from an early peak of 110.59 token/s to a latest reading of 9.17 token/s. In contrast, deepseek-v4-pro offers the most stable performance, maintaining an average of 83.7 token/s with a coefficient of variation of just 18.1%. The dataset contains 288 valid observations across all six models, achieving 100% coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 100.88 token/s, while minimax-m3 is the weakest at 43.57 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which exhibits a 60.7 percent coefficient of variation, frequent steep drops to near 6.83 token/s, and a negative trend of 23.8 percent. glm-5.2 also shows instability, plunging to 8.99 token/s at 09:10. The dataset has no missing-data limitation, as all six models maintain 48 valid samples out of 48 expected observations, yielding exactly 100.0 percent coverage.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 103.65 token/s, while minimax-m3 is the weakest at 42.17 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which exhibits a 56.7% coefficient of variation, frequent sharp drops to near 5 token/s, and a negative trend of 21.5%, ending at just 6.83 token/s. In contrast, minimax-m3 remains relatively stable with a 23.9% coefficient of variation. The dataset records 288 valid observations across six models, achieving 100.0% coverage. However, a missing-data limitation exists: the dataset entirely lacks concurrent request volume, preventing any correlation of throughput drops with concurrent load.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 107.14 token/s, while minimax-m3 is the weakest at 41.99 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which has a coefficient of variation of 57.7 percent and throughput swings from a minimum of 4.85 token/s to a maximum of 105.86 token/s. Despite this volatility, nemotron-3-ultra shows a positive trend of 19.6 percent, whereas minimax-m3 trends downward by 13.6 percent. The dataset includes 288 valid points across six models with 100.0 percent coverage, meaning there are no missing-data limitations to report for this period.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 107.04 token/s, while minimax-m3 is the weakest at 44.79 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which has a coefficient of variation of 55.5 percent and throughput swings from a low of 4.85 token/s to a high of 105.86 token/s. This erratic performance creates unpredictable capacity planning challenges. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage. However, a missing-data limitation exists because the dataset lacks observations outside this four-hour window, preventing any assessment of diurnal throughput patterns.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, gemma4:31b had the strongest average throughput at 103.52 token/s, while nemotron-3-ultra was weakest at 41.07 token/s. The most operationally significant volatility belongs to nemotron-3-ultra, which swung from a low of 4.85 token/s to a high of 105.86 token/s with a coefficient of variation of 69.5 percent, indicating highly unstable output capacity. In contrast, minimax-m3 delivered the most stable performance, averaging 45.43 token/s with a coefficient of variation of just 19.1 percent. Dataset coverage is complete with 288 valid points out of 288 expected, so there are no missing-data limitations affecting this specific analysis.

4-hour window · 288 points · 100.0% coverage

Over the four-hour period, gemma4:31b is the strongest model with an average throughput of 107.44 token/s, while minimax-m3 is the weakest at 44.64 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which has a coefficient of variation of 63.7 percent and throughput swings from a minimum of 5.26 token/s to a maximum of 94.13 token/s. In contrast, minimax-m3 shows the most stable performance with a standard deviation of 7.28 token/s. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage. However, a missing-data limitation exists because the dataset only captures five-minute intervals, completely omitting sub-interval latency drops or micro-bursts that would affect real-time user experience.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 111.39 token/s, while minimax-m3 is the weakest at 44.4 token/s. The most operationally significant volatility appears in nemotron-3-ultra, which has a coefficient of variation of 64.8 percent and throughput swings from a minimum of 5.26 token/s to a maximum of 109.5 token/s. deepseek-v4-flash also shows notable instability, dropping to 6.57 token/s around 02:05 before recovering. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitation over the expected 48 samples per model.

4-hour window · 288 points · 100.0% coverage

Over the four-hour window, gemma4:31b had the highest average throughput at 116.3 token/s, while nemotron-3-ultra was the weakest at 43.33 token/s. The most operationally significant volatility came from nemotron-3-ultra, which exhibited extreme instability with a coefficient of variation of 66.4 percent and sharp drops to a minimum of 5.69 token/s. Similarly, deepseek-v4-flash experienced a severe throughput drop to 6.57 token/s around 02:05 UTC. All six models recorded complete data with 48 out of 48 expected samples per model, resulting in 100.0 percent coverage and no missing-data limitations. Overall throughput declined toward the end of the period, with the latest measurements falling to 18.93 token/s for nemotron-3-ultra and 20.59 token/s for deepseek-v4-flash.

4-hour window · 276 points · 95.8% coverage

Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 116.87 token/s, while minimax-m3 is the weakest at 45.52 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits extreme instability with a coefficient of variation of 73.0 percent, swinging from a low of 3.13 token/s to a high of 150.01 token/s. Similarly, deepseek-v4-flash suffered a severe throughput drop to 6.57 token/s around 02:05 UTC. This analysis is constrained by a missing-data limitation; each model recorded 46 valid samples out of 48 expected, resulting in 95.8 percent coverage and leaving minor gaps in the rolling timeline.

4-hour window · 276 points · 95.8% coverage

Over the four-hour window, gemma4:31b is the strongest model with an average throughput of 115.28 token/s, while nemotron-3-ultra is the weakest at 36.26 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which has a coefficient of variation of 109.1% and swings wildly between a minimum of 1.93 token/s and a maximum of 150.01 token/s. This extreme instability contrasts with more consistent models like glm-5.2, which averages 110.99 token/s with a coefficient of variation of just 17.5%. A missing-data limitation affects this review, as each model has 46 valid samples out of 48 expected, resulting in 95.8% coverage and leaving minor gaps in the five-minute observation timeline.

4-hour window · 276 points · 95.8% coverage

Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 113.61 token/s, while nemotron-3-ultra is the weakest at 27.4 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which has a coefficient of variation of 142.2% and throughput swings from a minimum of 1.81 token/s to a maximum of 150.01 token/s. In contrast, minimax-m3 is the most stable with a coefficient of variation of 16.8% and an average of 48.75 token/s. A missing-data limitation affects this dataset, as each model has 46 valid samples out of 48 expected, resulting in 95.8% coverage.

4-hour window · 270 points · 93.8% coverage

Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 111.03 token/s, while nemotron-3-ultra is the weakest at 21.43 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which has a coefficient of variation of 153.0 percent and a minimum throughput of 1.81 token/s, despite a late spike to 150.01 token/s. In contrast, minimax-m3 is the most stable, maintaining an average of 49.06 token/s with a coefficient of variation of just 18.2 percent. This dataset is limited by missing observations; each model recorded 45 valid samples out of 48 expected, resulting in 93.8 percent coverage.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 107.52 token/s, while nemotron-3-ultra is the weakest at 28.58 token/s. The most operationally significant volatility comes from nemotron-3-ultra, which exhibits a coefficient of variation of 81.1 percent and a downward trend of 40.6 percent, dropping from a peak of 98.91 token/s to a final observation of just 3.34 token/s. Similarly, deepseek-v4-flash shows high instability with a coefficient of variation of 47.5 percent and frequent steep drops below 30 token/s. The dataset contains 288 valid observations across all models, achieving 100.0 percent coverage with no missing-data limitations.

4-hour window · 288 points · 100.0% coverage

Across the four-hour window, gemma4:31b had the strongest average output-token throughput at 95.83 token/s, while nemotron-3-ultra was weakest at 36.9 token/s. The most operationally significant volatility appeared in deepseek-v4-flash, which dropped sharply from 95.72 token/s at 19:55 to 12.87 token/s at 20:35 before recovering to 93.53 token/s by 20:55. This instability reflects its high coefficient of variation of 41.6 percent. The dataset includes 288 valid observations across six models, achieving 100.0 percent coverage with no missing-data limitation.