← All models

glm-5.3-flash

321.3B parameters FP8 1,048,576 token context completionthinkingtoolsvision
137.7
output tokens / second
160 tokens · 1.16s server · 1.30s request

Performance statistics

Calculated live from valid twenty-minute RRDtool samples.

Range Latest Average Lowest Median P50 P05 Highest Std dev Variability CV Trend Data coverage
Last 24 hours 137.8 109.9 35.3 105.4 57.4 196.4 35.7 32.5% -1.1% 72 samples 100.0%
Last 7 days 137.8 113.5 5.2 111.3 55.5 233.4 41.8 36.8% -4.6% 504 samples 100.0%
Last 30 days 137.8 91.7 3.0 78.2 44.4 233.4 41.7 45.5% +14.2% 2160 samples 100.0%

P05: 5% of twenty-minute throughput samples are at or below this value; it represents lower-tail performance. CV: standard deviation ÷ average; lower is more stable. Trend: second-half average versus first-half average. Data coverage: valid samples versus all expected twenty-minute slots in the selected range.

Last 24 hours

glm-5.3-flash throughput for Last 24 hours