- glm-5.3 is the strongest model by average throughput at 159.81 token/s (p95 of 182.91 token/s), while nemotron-3-ultra is the weakest at 26.08 token/s average, far below the next-lowest model, minimax-m3, at 68.89 token/s.
- The most operationally significant volatility is nemotron-3-ultra's step change around 22:25 UTC: it ran between 2.07 and 6.78 token/s for roughly the first two and a half hours, then jumped to peaks of 101.92 token/s, producing a 118.3% coefficient of variation.
- No missing-data limitation applies: all eight models recorded 48 of 48 expected samples, totaling 384 valid points at 100.0% coverage across the four-hour window.
Hourly performance insights
Summaries of rolling four-hour performance data
- glm-5.3 is the strongest model at 149.26 token/s average throughput, ahead of glm-5.2 at 135.11 token/s; nemotron-3-ultra is the weakest at 13.92 token/s average, with a minimum of 1.45 token/s.
- The most operationally significant volatility is nemotron-3-ultra, which ran near 2 to 6 token/s for most of the window, then spiked to 101.92 token/s at 22:35, driving its coefficient of variation to 185.1 percent.
- No missing-data limitation applies: all eight models report 48 of 48 samples at 100.0 percent coverage, and the dataset's 384 valid points match the expected total.
- glm-5.3 is the strongest model with average throughput of 149.09 token/s (p95 183.28 token/s), while nemotron-3-ultra is the weakest at 11.71 token/s average, peaking only once at 92.73 token/s.
- The most operationally significant event is nemotron-3-ultra's collapse: after 18:55 it drops from 92.73 token/s to 1.45 token/s by 19:40 and stays below 7 token/s through 22:00, a -78.9% trend with 160.7% coefficient of variation.
- No missing-data limitation applies: all eight models report 48 of 48 expected samples, 384 valid points total, and 100.0% coverage, so the four-hour window is fully represented.
- glm-5.3 is the strongest model by average throughput at 151.64 token/s (peak 186.34 token/s), while nemotron-3-ultra is the weakest at 16.46 token/s average (minimum 1.39 token/s).
- The most operationally significant volatility is nemotron-3-ultra's collapse: after 19:00 it stays between roughly 1.45 and 5.84 token/s through 21:00, versus an earlier peak of 92.73 token/s, producing a 135.2% coefficient of variation and a -89.4% trend.
- No missing-data limitation applies: all eight models report 48 of 48 expected samples with 100.0% coverage, and the dataset contains 384 valid points across the four-hour window.
- glm-5.3 is the strongest model with average throughput of 137.91 token/s (p95 183.4 token/s, max 186.34 token/s), while nemotron-3-ultra is the weakest at 19.77 token/s average, peaking at only 92.73 token/s.
- The most operationally significant volatility is nemotron-3-ultra: coefficient of variation 119.9%, with a sustained collapse from 19:05 to 20:00 UTC where readings stay between 1.45 and 5.84 token/s, versus its 92.73 token/s peak at 18:55.
- No missing-data limitation exists: all eight models report 48 of 48 samples with 100.0% coverage, and 384 valid points match the expected total, so the only constraint is the four-hour window itself.
- glm-5.2 is the strongest model with an average throughput of 124.14 token/s, while nemotron-3-ultra is the weakest at 31.39 token/s, roughly a quarter of the leader's rate.
- The most significant movement is glm-5.3's recovery: it started near 12.08 token/s at 15:05 UTC and climbed to a 186.34 token/s peak at 17:55, ending at 161.23 token/s. nemotron-3-ultra shows extreme volatility, with a coefficient of variation of 91.0% and swings between 1.35 and 109.36 token/s.
- No missing-data limitation exists: all eight models report 48 of 48 expected samples, 100.0% coverage, and 384 valid points across the four-hour window.
- glm-5.2 is the strongest model with an average throughput of 133.27 token/s (p95 176.08 token/s), while nemotron-3-ultra is the weakest at 30.40 token/s average, peaking at only 109.36 token/s.
- The most significant shift is glm-5.3, which climbed from single-digit rates near 6.57 token/s early in the window to a maximum of 186.34 token/s, a 255.5% trend with 74.1% coefficient of variation; nemotron-3-ultra was similarly unstable at 97.8% CV.
- No missing-data limitation exists: all eight models report 48 of 48 expected samples, 384 valid points total, and 100.0% coverage across the four-hour window.
- glm-5.2 delivered the highest average throughput at 150.74 token/s (p95 204.97 token/s), while nemotron-3-ultra was the weakest at 29.60 token/s average with a minimum of 0.92 token/s.
- glm-5.3 shows the most operationally significant volatility: throughput collapsed from 81.27 token/s at 14:00 to between 6.57 and 15.52 token/s from roughly 14:10 to 15:10, with a coefficient of variation of 69.3%, before recovering to 150.82 token/s by 17:00.
- No missing-data limitation applies: all eight models recorded 48 of 48 expected samples, giving 384 valid points and 100.0% coverage across the four-hour window.
- glm-5.2 is the strongest model at 160.44 token/s average throughput (peak 220.37 token/s), while nemotron-3-ultra is the weakest at 30.36 token/s average, ranging from 0.92 to 109.36 token/s across the window.
- glm-5.3 shows the most operationally significant volatility: it fell from 130.82 token/s at 12:50 to between 6.57 and 15.52 token/s from 14:10 to 15:15, producing a 74.5% coefficient of variation and a -53.3% trend, before partially recovering to 126.23 token/s at 15:45.
- glm-5.3 is missing 9 of 48 expected samples (81.2% coverage), with no observations from 12:05 to 12:45 UTC, so its early-window behavior is unknown; overall dataset coverage is 97.7% with 375 valid points.
- glm-5.2 is the strongest model at 166.2 token/s average throughput (p95 215.87 token/s); nemotron-3-ultra is the weakest at 24.42 token/s average, peaking at only 93.45 token/s and dropping to 0.92 token/s.
- glm-5.3 shows the most significant degradation: near 130 token/s at 12:50 UTC, it fell to 11.04 token/s at 14:10 and stayed between 6.57 and 14.67 token/s through 15:00, a 77.4% decline; nemotron-3-ultra is also highly volatile with 101.0% cv.
- glm-5.3 has only 27 of 48 expected samples (56.2% coverage) with no observations before 12:50 UTC, so its 64.68 token/s average is unreliable; overall coverage is 94.5% with 363 valid points.
- glm-5.2 is the strongest model at 162.28 token/s average throughput (p95 215.91 token/s), while nemotron-3-ultra is the weakest at 21.52 token/s average, never exceeding 63.67 token/s.
- nemotron-3-ultra shows the most operationally significant volatility, with a 97.6% coefficient of variation and swings between 1.15 and 63.67 token/s, including a 12:00–12:15 stretch stuck near 1–2.5 token/s; gemma4:31b also dipped to 24.09 token/s at 10:20.
- glm-5.3 has only 14 of 48 expected samples (29.2% coverage), all from 12:50 onward, so its 108.31 token/s average is not comparable to other models; overall dataset coverage is 89.3% across 343 valid points.
- Between 09:03 and 13:03 UTC, glm-5.2 was the strongest model at 159.75 token/s average output throughput, peaking at 223.09 token/s, while nemotron-3-ultra was weakest at 17.29 token/s average.
- The most operationally significant issue is nemotron-3-ultra's extreme volatility: a 113.1% coefficient of variation with swings between 1.15 and 63.01 token/s, including prolonged stretches near 2 token/s. By contrast, deepseek-v4-pro was the steadiest performer (10.9% CV, 109.68 token/s average).
- Overall coverage was 88.3% of 339 valid points; glm-5.3 reported only 3 of 48 expected samples (6.2% coverage), so its 96.95 token/s average is unreliable.
- Over the 08:56–12:56 UTC window, glm-5.2 was the strongest model, averaging 160.48 token/s (p95 210.27 token/s), while nemotron-3-ultra was the weakest at 16.45 token/s average.
- The most operationally significant issue is nemotron-3-ultra's extreme volatility: a 113.2% coefficient of variation, with five-minute readings swinging between 1.15 and 63.01 token/s, including prolonged stretches near 2 token/s. gemma4:31b also showed sharp swings (24.09–181.92 token/s) but trended upward 17.1%.
- Overall dataset coverage was 88.0% (338 valid points); glm-5.3 reported only 2 of 48 expected samples (4.2% coverage), so its 128.37 token/s average is unreliable and limits window-wide conclusions.
- Across the four-hour window, glm-5.2 delivered the strongest average throughput at 156.92 token/s, while nemotron-3-ultra was the weakest at 18.89 token/s.
- The most operationally significant volatility occurred in glm-5.2, which frequently swung from over 200 token/s down to brief severe drops near 16.66 token/s. In contrast, deepseek-v4-pro maintained the most stable performance, varying only between 52.41 and 131.48 token/s.
- Although dataset coverage is 100.0 percent across all 336 valid samples, the five-minute observation interval limits the ability to capture sub-minute latency spikes or exact drop durations.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 157.94 token/s, while nemotron-3-ultra was weakest at 17.67 token/s. Operationally, nemotron-3-ultra showed extreme volatility with a 109.6% coefficient of variation, dropping to a minimum of 1.48 token/s. Additionally, glm-5.2 and gemma4:31b exhibited sharp final-interval declines, falling to 16.66 token/s and 141.34 token/s, respectively. The dataset contains 336 valid points across seven models, achieving 100.0% coverage with no missing data limitations.
Across the four-hour window, glm-5.2 delivered the highest average throughput at 164.98 token/s, while nemotron-3-ultra was the weakest at 25.72 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which experienced a severe throughput collapse after 07:35 UTC, dropping from 40.20 token/s to 4.08 token/s and eventually hitting a minimum of 1.48 token/s. This model also exhibited extreme variability with a coefficient of variation of 95.7 percent. The dataset contains complete coverage with 336 valid points across all models, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 169.86 token/s, while nemotron-3-ultra was weakest at 35.49 token/s. Operationally, nemotron-3-ultra exhibited severe volatility with a 79.0% coefficient of variation and a 50.7% downward trend, dropping from a 105.57 token/s peak to just 1.55 token/s by 08:15. In contrast, minimax-m3 remained the most stable at 47.95 token/s average with a 15.3% coefficient of variation. The dataset is complete with 100.0% coverage and no missing-data limitations across all 336 valid observations.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 166.89 token/s, while nemotron-3-ultra was weakest at 39.12 token/s. Operationally, nemotron-3-ultra showed extreme volatility with a 71.9% coefficient of variation, dropping from a peak of 105.57 token/s to a low of 2.72 token/s. In contrast, minimax-m3 was the most stable model with a 17.1% coefficient of variation and an average of 49.09 token/s. The dataset contains 336 valid observations across seven models with 100.0% coverage, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 169.82 token/s, while nemotron-3-ultra was weakest at 47.52 token/s. Operationally, nemotron-3-ultra exhibited extreme volatility, swinging from a low of 4.31 to a high of 116.48 token/s, yielding a coefficient of variation of 59.8 percent. This instability contrasts sharply with the steady performance of glm-5.3-flash, which maintained a much tighter range between 50.51 and 102.02 token/s. The dataset is complete with 100.0 percent coverage across all 336 valid points, meaning there are no missing-data limitations to account for in this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 172.07 token/s, while nemotron-3-ultra was weakest at 48.92 token/s. The most operationally significant volatility is nemotron-3-ultra's extreme instability, swinging from a high of 120.56 token/s down to 4.31 token/s, yielding a 64.9 percent coefficient of variation and a 23.1 percent downward trend. In contrast, glm-5.3-flash remained the most stable, averaging 87.03 token/s with a 9.7 percent coefficient of variation. Although dataset coverage is 100.0 percent with 336 valid samples, the four-hour duration limits broader operational conclusions.
Across the four-hour window, glm-5.2 delivered the highest average throughput at 163.79 token/s, while nemotron-3-ultra was the weakest at 48.87 token/s. Operationally, nemotron-3-ultra exhibited extreme volatility with a 59.7% coefficient of variation, swinging from a maximum of 120.56 token/s down to 4.31 token/s. Additionally, deepseek-v4-pro showed a notable downward trend, dropping 11.5% over the period to a low of 32.43 token/s. The dataset contains 336 valid observations across seven models with 100.0% coverage, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 156.45 token/s, while nemotron-3-ultra was weakest at 40.18 token/s. The most operationally significant volatility belongs to nemotron-3-ultra, which dropped to 0.9 token/s and showed a 82.2% coefficient of variation, indicating highly unstable performance. In contrast, glm-5.3-flash was the most stable model, averaging 85.48 token/s with a low 9.8% coefficient of variation. A missing-data limitation affects this analysis: nemotron-3-ultra has 47 valid samples instead of the expected 48, resulting in 97.9% coverage. All other models maintained complete 100% coverage across the 335 total valid observations.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 156.43 token/s, while nemotron-3-ultra was weakest at 40.25 token/s. Operationally, nemotron-3-ultra exhibited extreme volatility, dropping to 0.9 token/s around 00:50 UTC before recovering to 145.81 token/s earlier at 23:55 UTC. This volatility reflects a coefficient of variation of 86.7 percent. A missing-data limitation affects nemotron-3-ultra, which has 47 valid samples instead of the expected 48, resulting in 97.9 percent coverage. In contrast, glm-5.3-flash maintained the most stable performance with a low 11.0 percent coefficient of variation and an average of 84.07 token/s.
Across the four-hour window, glm-5.2 delivered the highest average throughput at 156.28 token/s, while nemotron-3-ultra was the weakest at 34.85 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which exhibited extreme instability with a coefficient of variation of 91.2% and a sharp downward trend of -45.9%, plummeting from a peak of 145.81 token/s to near-zero levels around 00:50. This dataset contains a missing-data limitation: nemotron-3-ultra has only 47 valid samples out of an expected 48, missing one five-minute observation entirely. All other models maintained complete 100.0% coverage.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 157.82 token/s, while nemotron-3-ultra was weakest at 35.96 token/s. Operationally, nemotron-3-ultra showed extreme volatility with a 94.2% coefficient of variation and a severe downward trend, dropping from 145.81 token/s to near-zero levels around 00:10. In contrast, glm-5.3-flash maintained the most stable performance with a 12.5% coefficient of variation. The dataset has a minor missing-data limitation: nemotron-3-ultra recorded 47 samples instead of the expected 48, missing one observation, though overall dataset coverage remained high at 99.7%.
Over the four-hour window, glm-5.2 delivered the highest average throughput at 165.6 token/s, while minimax-m3 was the weakest at 47.71 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which exhibited extreme swings from a low of 4.89 token/s to a high of 145.81 token/s, resulting in a coefficient of variation of 70.0 percent. Several other models, including glm-5.2 and deepseek-v4-flash, also experienced periodic sharp throughput drops below 25 token/s. The dataset contains complete coverage with 336 valid points across all seven models, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 159.66 token/s, while nemotron-3-ultra was weakest at 38.3 token/s. Operationally, glm-5.2 showed high volatility with a 30.7% coefficient of variation, including severe drops to 25.73 token/s. Nemotron-3-ultra exhibited even greater instability with a 68.0% coefficient of variation, ranging from 4.89 to 135.27 token/s. The dataset contains 336 valid points across seven models with 100.0% coverage, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 143.96 token/s, while nemotron-3-ultra was weakest at 31.39 token/s. The most operationally significant volatility occurred in glm-5.2, which fluctuated heavily between 10.73 and 217.13 token/s despite its high average. Nemotron-3-ultra also showed extreme instability, starting near 1.8 token/s before spiking to 135.27 token/s at the end. The dataset contains 336 valid points across seven models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 144.6 token/s, while nemotron-3-ultra was weakest at 20.66 token/s. The most operationally significant volatility occurred in glm-5.2, which fluctuated from a low of 10.73 token/s to a high of 217.13 token/s, and in nemotron-3-ultra, which showed extreme instability with a coefficient of variation of 104.2 percent. Deepseek-v4-flash exhibited a notable throughput drop late in the window, falling to 46.28 token/s at 20:25. The dataset contains 336 valid points across seven models, achieving 100.0 percent coverage with no missing-data limitation.
Across the four-hour window, glm-5.2 had the strongest average throughput at 152.29 token/s, while nemotron-3-ultra was weakest at 13.5 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from 203.35 token/s at 17:50 to 10.73 token/s at 18:10 before recovering. Nemotron-3-ultra also exhibited extreme instability, fluctuating between 0.6 and 75.07 token/s with a coefficient of variation of 134.8 percent. A missing-data limitation affects nemotron-3-ultra, which recorded only 45 valid samples out of 48 expected, resulting in 93.8 percent coverage and leaving three five-minute intervals unobserved.
Across the four-hour window, glm-5.2 delivered the highest average throughput at 158.21 token/s, while nemotron-3-ultra was the weakest at 4.7 token/s. The most operationally significant volatility occurred in glm-5.2, which experienced severe latency spikes dropping throughput to 10.73 token/s at 18:10 before recovering. Overall dataset coverage is 98.8%, but this masks a missing-data limitation for nemotron-3-ultra, which has only 44 valid samples out of 48 expected, restricting visibility into its true performance baseline.
Over the four-hour window, glm-5.2 delivered the strongest average throughput at 164.14 token/s, while nemotron-3-ultra was the weakest at 2.85 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from 213.81 token/s at 17:10 to 25.55 token/s at 17:15 before recovering. Deepseek-v4-flash showed a notable upward trend, increasing by 10.0 percent to an average of 101.5 token/s. A key limitation is missing data: nemotron-3-ultra has only 44 valid samples out of 48 expected, resulting in 91.7 percent coverage, which leaves gaps in its throughput profile.
Across the four-hour window, glm-5.2 delivered the highest average throughput at 163.88 token/s, while nemotron-3-ultra was the weakest at 1.64 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which spiked from sub-1 token/s baselines to 6.35 token/s after 16:30 UTC. Gemma4:31b also showed notable instability, swinging between 24.51 and 136.67 token/s. A missing-data limitation affects this analysis: nemotron-3-ultra recorded only 43 of 48 expected samples, dropping five observations between 16:10 and 16:30 UTC, meaning its recent throughput spike lacks continuous context.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 141.69 token/s, while nemotron-3-ultra was weakest at 0.91 token/s. The most operationally significant volatility was a severe throughput drop around 13:20 UTC, where gemma4:31b fell to 24.51 token/s, glm-5.2 dropped to 42.09 token/s, and glm-5.3-flash reached 24.24 token/s. Despite this drop, glm-5.2 showed a strong upward trend of 39.8 percent, ending at 183.02 token/s. The dataset has a missing-data limitation: nemotron-3-ultra recorded only 44 samples instead of the expected 48, resulting in 91.7 percent coverage, which limits visibility into its exact performance dips.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 121.03 token/s, while nemotron-3-ultra was the weakest at 2.83 token/s. The most operationally significant volatility is glm-5.2's sharp 64.4 percent upward trend, which included a peak of 212.94 token/s despite a high coefficient of variation of 37.1 percent. In contrast, deepseek-v4-pro offered the most stable performance, averaging 102.65 token/s with a low coefficient of variation of 12.9 percent. A missing-data limitation affects this analysis: nemotron-3-ultra captured only 44 of 48 expected samples, dropping to 91.7 percent coverage, which limits the reliability of its calculated averages.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 107.11 token/s, while nemotron-3-ultra was weakest at 8.72 token/s. Nemotron-3-ultra also showed the most operationally significant volatility, with a 161.5% coefficient of variation and a 94.5% downward trend, dropping from a 64.81 token/s peak to 1.17 token/s. In contrast, deepseek-v4-pro offered the most stable performance, averaging 104.33 token/s with a 14.1% coefficient of variation. A missing-data limitation affects nemotron-3-ultra, which recorded only 45 valid samples instead of the expected 48, resulting in 93.8% coverage.
Across the four-hour window, deepseek-v4-pro delivered the strongest average throughput at 106.16 token/s, while nemotron-3-ultra was weakest at 11.18 token/s. The most operationally significant volatility is nemotron-3-ultra's severe degradation, trending down 69.9 percent to a latest throughput of 1.05 token/s. Conversely, deepseek-v4-flash showed a positive 15.5 percent trend, peaking at 154.18 token/s. A missing-data limitation affects this analysis: nemotron-3-ultra recorded only 46 valid samples instead of 48, dropping its coverage to 95.8 percent, whereas the other six models maintained complete 48-sample coverage.
Across the four-hour window, deepseek-v4-pro had the strongest average output-token throughput at 105.69 token/s, while nemotron-3-ultra was the weakest at 13.14 token/s. The most operationally significant volatility came from nemotron-3-ultra, which exhibited a 95.0% coefficient of variation and dropped to a minimum of 0.77 token/s. In contrast, deepseek-v4-pro maintained the most stable performance with a 14.0% coefficient of variation. Overall dataset coverage was 99.7% across 335 valid points, but a missing-data limitation exists for nemotron-3-ultra, which recorded only 47 samples instead of the expected 48, leaving a gap in its observations.
Across the four-hour observation period, deepseek-v4-pro achieved the strongest average throughput at 106.99 token/s, while nemotron-3-ultra was the weakest at 13.72 token/s. The most operationally significant volatility occurred in nemotron-3-ultra, which exhibited a coefficient of variation of 78.4 percent and a sharp throughput spike to 64.81 token/s at 10:55 UTC after hovering near 10 token/s for most of the window. In contrast, deepseek-v4-pro maintained the most stable performance with a coefficient of variation of just 15.4 percent. The dataset contains complete coverage with 336 valid points across all seven models, so there are no missing-data limitations affecting this specific interval.
Across the four-hour observation window, deepseek-v4-pro achieved the highest average throughput at 105.82 token/s, while nemotron-3-ultra was the weakest at 11.38 token/s. The most operationally significant volatility occurred within nemotron-3-ultra, which exhibited extreme instability with a coefficient of variation of 69.5 percent and throughput frequently dropping below 10 token/s, despite starting at 45.46 token/s. In contrast, deepseek-v4-pro maintained the most stable performance with a coefficient of variation of just 16.1 percent. The dataset contains 336 valid observations across seven models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this analysis.
Across the four-hour window, deepseek-v4-pro had the strongest average throughput at 106.06 token/s, while nemotron-3-ultra was weakest at 16.65 token/s. The most operationally significant volatility appeared in nemotron-3-ultra, which dropped from a 78.58 token/s peak to a 2.46 token/s minimum with a 102.1 percent coefficient of variation, alongside a severe downward trend of 54.8 percent. In contrast, deepseek-v4-pro remained stable with a 16.4 percent coefficient of variation. glm-5.2 also showed notable instability, falling to 21.35 token/s at 06:10. The dataset contains 336 valid observations with 100.0 percent coverage, so no missing-data limitation affects this specific analysis.
Across the four-hour window, gemma4:31b delivered the strongest average throughput at 115.32 token/s, while nemotron-3-ultra was the weakest at 22.98 token/s. The most operationally significant volatility came from nemotron-3-ultra, which dropped 62.4 percent over the period to a latest throughput of 8.19 token/s, alongside a coefficient of variation of 87.6 percent. deepseek-v4-flash showed a contrasting upward trend of 13.4 percent but remained highly volatile, swinging between 13.7 and 146.06 token/s. The dataset includes 336 valid observations across seven models with 100.0 percent coverage, so there are no missing-data limitations affecting this analysis.
Across the four-hour window, gemma4:31b had the strongest average throughput at 115.75 token/s, while nemotron-3-ultra was weakest at 31.28 token/s. The most operationally significant volatility appeared in deepseek-v4-flash, which dropped from 149.62 token/s to a minimum of 13.7 token/s before recovering, yielding a coefficient of variation of 52.2 percent. Nemotron-3-ultra also showed severe instability, trending down by 42.1 percent and ending at 7.15 token/s. In contrast, deepseek-v4-pro remained the most stable, averaging 105.85 token/s with a coefficient of variation of just 15.8 percent. The dataset includes 336 valid observations across seven models, achieving 100.0 percent coverage with no missing-data limitations.
Across the four-hour observation window, gemma4:31b achieved the highest average throughput at 124.69 token/s, while nemotron-3-ultra was the weakest at 34.62 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which experienced a severe throughput collapse from 149.62 token/s down to 13.7 token/s between 03:20 and 04:05, reflecting a 51.6 coefficient of variation. Nemotron-3-ultra also showed extreme instability, dropping to a minimum of 3.69 token/s. The dataset contains 336 valid observations across seven models with 100.0 percent coverage, so no missing-data limitations affect this specific analysis window.
Across the four-hour window, gemma4:31b is the strongest model with an average throughput of 126.65 token/s, while nemotron-3-ultra is the weakest at 32.78 token/s. The most operationally significant volatility appears in deepseek-v4-flash, which dropped sharply from 132.42 token/s at 02:35 to 13.7 token/s at 04:05 before recovering to 140.47 token/s at 04:45. Nemotron-3-ultra also showed extreme instability, spiking to 94.97 token/s at 03:05 after a low of 3.69 token/s at 02:45. The dataset includes 336 valid observations across seven models. Although coverage is complete at 100 percent, the five-minute sampling interval limits the ability to detect sub-minute latency spikes or brief micro-stalls.
Across the four-hour observation window, gemma4:31b delivered the strongest average output-token throughput at 118.25 token/s, while nemotron-3-ultra was the weakest at 30.27 token/s. The most operationally significant volatility occurred in deepseek-v4-flash, which maintained an average of 89.54 token/s before suffering a severe throughput collapse after 03:30 UTC, dropping to a minimum of 18.16 token/s. Nemotron-3-ultra also exhibited extreme instability, ranging from 3.59 to 94.97 token/s with a coefficient of variation of 70.3 percent. The dataset contains 336 valid observations across seven models with 100.0 percent coverage, meaning there are no missing-data limitations affecting this analysis.
Across the four-hour window, glm-5.2 delivered the highest average throughput at 123.04 token/s, while nemotron-3-ultra was the weakest at 29.26 token/s. The most operationally significant volatility came from glm-5.3-flash, which exhibited extreme swings including a minimum of 17.49 token/s and a maximum of 152.6 token/s, reflecting a 39.3% coefficient of variation. Additionally, nemotron-3-ultra experienced severe periodic drops, plunging to 3.59 token/s. The dataset contains 336 valid observations across seven models, achieving 100.0% coverage. However, a missing-data limitation exists: the dataset lacks any concurrent request volume or concurrency metrics, making it impossible to determine if these throughput fluctuations were caused by varying load or internal system instability.
Across the four-hour window, glm-5.2 delivered the strongest average throughput at 129.67 token/s, while nemotron-3-ultra was the weakest at 26.97 token/s. The most operationally significant volatility comes from glm-5.3-flash, which spiked from a 17.49 token/s minimum to a 149.62 token/s maximum, yielding a 42.5% coefficient of variation and a 55.7% upward trend. Nemotron-3-ultra also showed extreme instability, dropping as low as 1.73 token/s with a 90.6% coefficient of variation. The dataset contains 336 valid observations across seven models with 100.0% coverage, so there are no missing-data limitations affecting this analysis.
Across the four-hour observation window, glm-5.2 delivered the strongest average output-token throughput at 127.97 token/s, while nemotron-3-ultra was the weakest at 25.45 token/s. The most operationally significant volatility came from glm-5.3-flash, which dropped sharply from 154.48 token/s to 17.49 token/s, reflecting a 44.6% coefficient of variation. Similarly, nemotron-3-ultra exhibited extreme instability, fluctuating between 1.73 token/s and 89.2 token/s. The dataset contains 336 valid observations across seven models, achieving 100.0% coverage with no missing-data limitations. This complete dataset allows for reliable operational assessments of model throughput without gaps.
Across the four-hour observation window, gemma4:31b delivered the strongest average output-token throughput at 124.01 token/s, while nemotron-3-ultra was the weakest at 27.11 token/s. The most operationally significant volatility came from glm-5.3-flash, which experienced a 51.7 percent downward trend, dropping from early peaks above 150 token/s to a low of 17.49 token/s. Nemotron-3-ultra also showed extreme instability, oscillating between 0.99 and 89.2 token/s. The dataset contains 336 valid observations across seven models with 100.0 percent coverage, so there are no missing-data limitations affecting this specific period.