Ollama Cloud Performance

Output-token throughput benchmarks every 5 minutes

6/6 models online
gemma4:31b
40.9
output tokens / second
125 tokens · 3.2s request
minimax-m3
44.2
output tokens / second
160 tokens · 3.8s request
glm-5.2
63.4
output tokens / second
160 tokens · 2.7s request
deepseek-v4-pro
92.2
output tokens / second
160 tokens · 1.9s request
deepseek-v4-flash
106.7
output tokens / second
160 tokens · 1.7s request
nemotron-3-ultra
50.9
output tokens / second
160 tokens · 3.3s request

Hourly performance insights

GLM-5.2 analysis of the rolling four-hour, five-minute dataset

AI-generated · hourly
4-hour window · 288 points · 100.0% coverage

Across the four-hour window, deepseek-v4-flash achieved the strongest average throughput at 103.84 token/s, while nemotron-3-ultra was the weakest at 45.14 token/s. The most operationally significant volatility occurred in glm-5.2, which dropped sharply from a peak of 216.98 token/s down to 13.59 token/s at 04:30, reflecting a coefficient of variation of 53.4 percent. This dataset contains 288 valid observations across six models, corresponding to exactly 100.0 percent coverage of the 48 expected samples per model, meaning there are no missing-data limitations to report.

View full insight feed →

Last 24 hours

Token throughput for Last 24 hours