Ollama Cloud Performance

Output-token throughput benchmarks every 20 minutes

8/8 models online
gemma4:31b
155.1
output tokens / second
123 tokens · 1.0s request
minimax-m3
39.9
output tokens / second
160 tokens · 4.2s request
glm-5.2
98.8
output tokens / second
160 tokens · 1.9s request
glm-5.3
80.5
output tokens / second
160 tokens · 2.3s request
glm-5.3-flash
151.4
output tokens / second
160 tokens · 1.4s request
deepseek-v4-pro
102.9
output tokens / second
160 tokens · 1.9s request
deepseek-v4.1-flash
164.9
output tokens / second
160 tokens · 1.1s request
nemotron-3-ultra
52.8
output tokens / second
160 tokens · 3.4s request

Hourly performance insights

Analysis of the rolling four-hour, twenty-minute dataset

AI-generated · hourly
4-hour window · 96 points · 100.0% coverage
  • deepseek-v4.1-flash is the strongest model at 186.75 token/s average throughput (peak 235.93 token/s), while nemotron-3-ultra is the weakest at 22.48 token/s average (minimum 3.15 token/s).
  • glm-5.2 shows the most extreme volatility, ranging from 10.78 to 216.7 token/s with a 94.6% coefficient of variation; gemma4:31b also swung from 146.13 down to 43.95 token/s late in the window.
  • No missing-data limitation exists: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the four-hour window.
View full insight feed →

Last 24 hours

Token throughput for Last 24 hours