{"summaries":[{"coverage_pct":100.0,"eval_count":766,"generated_at":"2026-10-08T04:03:09.600812+00:00","generated_label":"Oct 08, 2026 \u00b7 04:03 UTC","id":1210,"model":"glm-5.3","period_end":"2026-10-08T04:03:01+00:00","period_label":"Oct 08, 2026 \u00b7 00:03 UTC to Oct 08, 2026 \u00b7 04:03 UTC","period_start":"2026-10-08T00:03:01+00:00","prompt_eval_count":2247,"summary":"- deepseek-v4.1-flash is the strongest model at 174.07 token/s average throughput, while nemotron-3-ultra is the weakest at 26.30 token/s; gemma4:31b and glm-5.3-flash follow at 111.48 and 113.90 token/s respectively.\n- glm-5.2 shows the highest volatility with a 76.8% coefficient of variation, ranging from 12.80 to 169.40 token/s, including a spike to its maximum in the final 04:00 observation; minimax-m3 declined 30.2% overall, ending at 17.50 token/s.\n- No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 174.07 token/s average throughput, while nemotron-3-ultra is the weakest at 26.30 token/s; gemma4:31b and glm-5.3-flash follow at 111.48 and 113.90 token/s respectively.\n- glm-5.2 shows the highest volatility with a 76.8% coefficient of variation, ranging from 12.80 to 169.40 token/s, including a spike to its maximum in the final 04:00 observation; minimax-m3 declined 30.2% overall, ending at 17.50 token/s.\n- No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 174.07 token/s average throughput, while nemotron-3-ultra is the weakest at 26.30 token/s; gemma4:31b and glm-5.3-flash follow at 111.48 and 113.90 token/s respectively.","glm-5.2 shows the highest volatility with a 76.8% coefficient of variation, ranging from 12.80 to 169.40 token/s, including a spike to its maximum in the final 04:00 observation; minimax-m3 declined 30.2% overall, ending at 17.50 token/s.","No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 174.07 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 26.30 \u8bcd\u5143/\u79d2\uff1bgemma4:31b \u548c glm-5.3-flash \u5206\u522b\u4ee5 111.48 \u548c 113.90 \u8bcd\u5143/\u79d2\u7d27\u968f\u5176\u540e\u3002\n- glm-5.2 \u6ce2\u52a8\u6027\u6700\u9ad8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 76.8%\uff0c\u8303\u56f4\u4ece 12.80 \u5230 169.40 \u8bcd\u5143/\u79d2\uff0c\u5305\u62ec\u5728\u6700\u540e 04:00 \u7684\u89c2\u6d4b\u4e2d\u98d9\u5347\u81f3\u5176\u6700\u5927\u503c\uff1bminimax-m3 \u6574\u4f53\u4e0b\u964d 30.2%\uff0c\u6700\u7ec8\u4e3a 17.50 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u7f3a\u5931\u6570\u636e\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u8bb0\u5f55\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u5171\u4ea7\u751f 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":5737340734,"translated_at":"2026-10-08T04:03:09.600812+00:00","translation_eval_count":210,"translation_model":"glm-5.3","translation_prompt_eval_count":350,"translation_total_duration_ns":2186515262,"translation_wall_seconds":2.332,"valid_point_count":96,"wall_seconds":5.915},{"coverage_pct":100.0,"eval_count":541,"generated_at":"2026-10-08T02:03:05.295907+00:00","generated_label":"Oct 08, 2026 \u00b7 02:03 UTC","id":1209,"model":"glm-5.3","period_end":"2026-10-08T02:03:01+00:00","period_label":"Oct 07, 2026 \u00b7 22:03 UTC to Oct 08, 2026 \u00b7 02:03 UTC","period_start":"2026-10-07T22:03:01+00:00","prompt_eval_count":2245,"summary":"- deepseek-v4.1-flash is the strongest model at 178.6 token/s average throughput, peaking at 225.67 token/s; nemotron-3-ultra is the weakest at 31.92 token/s average, never exceeding 68.52 token/s.\n- The most operationally significant volatility is glm-5.3-flash swinging from 19.95 token/s at 23:20 to 220.72 token/s at 23:40, while glm-5.3 shows the steepest decline at -24.6% trend, dipping to 60.12 token/s at 00:00 before recovering to 163.28 token/s by 02:00.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 178.6 token/s average throughput, peaking at 225.67 token/s; nemotron-3-ultra is the weakest at 31.92 token/s average, never exceeding 68.52 token/s.\n- The most operationally significant volatility is glm-5.3-flash swinging from 19.95 token/s at 23:20 to 220.72 token/s at 23:40, while glm-5.3 shows the steepest decline at -24.6% trend, dipping to 60.12 token/s at 00:00 before recovering to 163.28 token/s by 02:00.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 178.6 token/s average throughput, peaking at 225.67 token/s; nemotron-3-ultra is the weakest at 31.92 token/s average, never exceeding 68.52 token/s.","The most operationally significant volatility is glm-5.3-flash swinging from 19.95 token/s at 23:20 to 220.72 token/s at 23:40, while glm-5.3 shows the steepest decline at -24.6% trend, dipping to 60.12 token/s at 00:00 before recovering to 163.28 token/s by 02:00.","No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 178.6 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 225.67 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 31.92 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 68.52 \u8bcd\u5143/\u79d2\u3002\n- \u8fd0\u8425\u5c42\u9762\u6700\u663e\u8457\u7684\u6ce2\u52a8\u662f glm-5.3-flash \u4ece 23:20 \u7684 19.95 \u8bcd\u5143/\u79d2\u6446\u52a8\u81f3 23:40 \u7684 220.72 \u8bcd\u5143/\u79d2\uff0c\u800c glm-5.3 \u5448\u73b0\u6700\u9661\u5ced\u7684\u4e0b\u964d\u8d8b\u52bf\uff0c\u4e3a -24.6%\uff0c\u5728 00:00 \u8dcc\u81f3 60.12 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u5728 02:00 \u524d\u56de\u5347\u81f3 163.28 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0cvalid_point_count \u4e3a 96\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":2246256163,"translated_at":"2026-10-08T02:03:05.295907+00:00","translation_eval_count":217,"translation_model":"glm-5.3","translation_prompt_eval_count":360,"translation_total_duration_ns":1246789025,"translation_wall_seconds":1.389,"valid_point_count":96,"wall_seconds":2.591},{"coverage_pct":100.0,"eval_count":894,"generated_at":"2026-10-08T01:03:07.113335+00:00","generated_label":"Oct 08, 2026 \u00b7 01:03 UTC","id":1208,"model":"glm-5.3","period_end":"2026-10-08T01:03:01+00:00","period_label":"Oct 07, 2026 \u00b7 21:03 UTC to Oct 08, 2026 \u00b7 01:03 UTC","period_start":"2026-10-07T21:03:01+00:00","prompt_eval_count":2245,"summary":"- deepseek-v4.1-flash is the strongest model at 162.01 token/s average throughput, peaking at 225.67 token/s; nemotron-3-ultra is the weakest at 24.79 token/s average, never exceeding 64.44 token/s.\n- Throughput declined across most models over the window: glm-5.3 fell 27.2% and glm-5.2 fell 25.4%, while deepseek-v4.1-flash rose 38.8%. Volatility was highest for nemotron-3-ultra (65.5% coefficient of variation) and minimax-m3 (51.8%), versus deepseek-v4-pro's steadier 17.4%.\n- No missing-data limitation applies: all 96 expected points are valid, with 100.0% coverage and 12 of 12 samples present for each of the eight models.","summary_en":"- deepseek-v4.1-flash is the strongest model at 162.01 token/s average throughput, peaking at 225.67 token/s; nemotron-3-ultra is the weakest at 24.79 token/s average, never exceeding 64.44 token/s.\n- Throughput declined across most models over the window: glm-5.3 fell 27.2% and glm-5.2 fell 25.4%, while deepseek-v4.1-flash rose 38.8%. Volatility was highest for nemotron-3-ultra (65.5% coefficient of variation) and minimax-m3 (51.8%), versus deepseek-v4-pro's steadier 17.4%.\n- No missing-data limitation applies: all 96 expected points are valid, with 100.0% coverage and 12 of 12 samples present for each of the eight models.","summary_items":["deepseek-v4.1-flash is the strongest model at 162.01 token/s average throughput, peaking at 225.67 token/s; nemotron-3-ultra is the weakest at 24.79 token/s average, never exceeding 64.44 token/s.","Throughput declined across most models over the window: glm-5.3 fell 27.2% and glm-5.2 fell 25.4%, while deepseek-v4.1-flash rose 38.8%. Volatility was highest for nemotron-3-ultra (65.5% coefficient of variation) and minimax-m3 (51.8%), versus deepseek-v4-pro's steadier 17.4%.","No missing-data limitation applies: all 96 expected points are valid, with 100.0% coverage and 12 of 12 samples present for each of the eight models."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 162.01 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 225.67 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 24.79 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 64.44 \u8bcd\u5143/\u79d2\u3002\n- \u5728\u89c2\u6d4b\u65f6\u6bb5\u5185\uff0c\u5927\u591a\u6570\u6a21\u578b\u7684\u541e\u5410\u91cf\u6709\u6240\u4e0b\u964d\uff1aglm-5.3 \u4e0b\u964d 27.2%\uff0cglm-5.2 \u4e0b\u964d 25.4%\uff0c\u800c deepseek-v4.1-flash \u4e0a\u5347 38.8%\u3002\u6ce2\u52a8\u6027\u6700\u9ad8\u7684\u662f nemotron-3-ultra\uff08\u53d8\u5f02\u7cfb\u6570 65.5%\uff09\u548c minimax-m3\uff0851.8%\uff09\uff0c\u76f8\u6bd4\u4e4b\u4e0b deepseek-v4-pro \u66f4\u4e3a\u7a33\u5b9a\uff0c\u4e3a 17.4%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8 96 \u4e2a\u9884\u671f\u6570\u636e\u70b9\u5747\u6709\u6548\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u516b\u4e2a\u6a21\u578b\u4e2d\u6bcf\u4e2a\u6a21\u578b\u90fd\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684\u5168\u90e8 12 \u4e2a\u3002","total_duration_ns":3997543436,"translated_at":"2026-10-08T01:03:07.113335+00:00","translation_eval_count":217,"translation_model":"glm-5.3","translation_prompt_eval_count":363,"translation_total_duration_ns":1062624249,"translation_wall_seconds":1.204,"valid_point_count":96,"wall_seconds":4.32},{"coverage_pct":100.0,"eval_count":290,"generated_at":"2026-10-08T00:03:07.093611+00:00","generated_label":"Oct 08, 2026 \u00b7 00:03 UTC","id":1207,"model":"glm-5.3","period_end":"2026-10-08T00:03:01+00:00","period_label":"Oct 07, 2026 \u00b7 20:03 UTC to Oct 08, 2026 \u00b7 00:03 UTC","period_start":"2026-10-07T20:03:01+00:00","prompt_eval_count":2246,"summary":"- deepseek-v4.1-flash is the strongest model at 166.76 token/s average throughput, while nemotron-3-ultra is the weakest at 27.55 token/s average, roughly six times slower.\n- The most operationally significant volatility is minimax-m3, with a 62.7% coefficient of variation, swinging between 118.54 token/s at 21:20 and 3.84 token/s at 00:00; glm-5.2 also shows the steepest decline at -26.7% trend, ending at 36.34 token/s.\n- No missing-data limitation exists: all 8 models recorded 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 166.76 token/s average throughput, while nemotron-3-ultra is the weakest at 27.55 token/s average, roughly six times slower.\n- The most operationally significant volatility is minimax-m3, with a 62.7% coefficient of variation, swinging between 118.54 token/s at 21:20 and 3.84 token/s at 00:00; glm-5.2 also shows the steepest decline at -26.7% trend, ending at 36.34 token/s.\n- No missing-data limitation exists: all 8 models recorded 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 166.76 token/s average throughput, while nemotron-3-ultra is the weakest at 27.55 token/s average, roughly six times slower.","The most operationally significant volatility is minimax-m3, with a 62.7% coefficient of variation, swinging between 118.54 token/s at 21:20 and 3.84 token/s at 00:00; glm-5.2 also shows the steepest decline at -26.7% trend, ending at 36.34 token/s.","No missing-data limitation exists: all 8 models recorded 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 166.76 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 27.55 \u8bcd\u5143/\u79d2\uff0c\u901f\u5ea6\u5927\u7ea6\u6162\u516d\u500d\u3002\n- \u8fd0\u8425\u5c42\u9762\u6700\u663e\u8457\u7684\u6ce2\u52a8\u6765\u81ea minimax-m3\uff0c\u5176\u53d8\u5f02\u7cfb\u6570\u4e3a 62.7%\uff0c\u5728 21:20 \u7684 118.54 \u8bcd\u5143/\u79d2\u4e0e 00:00 \u7684 3.84 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u5927\u5e45\u6ce2\u52a8\uff1bglm-5.2 \u4e5f\u8868\u73b0\u51fa\u6700\u9661\u5ced\u7684\u4e0b\u964d\u8d8b\u52bf\uff0c\u4e3a -26.7%\uff0c\u6700\u7ec8\u6536\u4e8e 36.34 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8 8 \u4e2a\u6a21\u578b\u5747\u8bb0\u5f55\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u5171\u4ea7\u751f 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":2481274438,"translated_at":"2026-10-08T00:03:07.093611+00:00","translation_eval_count":202,"translation_model":"glm-5.3","translation_prompt_eval_count":337,"translation_total_duration_ns":1979821093,"translation_wall_seconds":2.299,"valid_point_count":96,"wall_seconds":2.81},{"coverage_pct":100.0,"eval_count":588,"generated_at":"2026-10-07T23:03:05.688098+00:00","generated_label":"Oct 07, 2026 \u00b7 23:03 UTC","id":1206,"model":"glm-5.3","period_end":"2026-10-07T23:03:01+00:00","period_label":"Oct 07, 2026 \u00b7 19:03 UTC to Oct 07, 2026 \u00b7 23:03 UTC","period_start":"2026-10-07T19:03:01+00:00","prompt_eval_count":2246,"summary":"- deepseek-v4.1-flash is the strongest model at 168.93 token/s average throughput, peaking at 230.47 token/s at 20:20 UTC; nemotron-3-ultra is the weakest at 24.86 token/s average, ending at 6.91 token/s at 23:00 UTC.\n- The most operationally significant volatility is deepseek-v4.1-flash collapsing to 15.72 token/s at 22:00 UTC from 215.07 token/s at 21:40, driving a 36.3% coefficient of variation; nemotron-3-ultra also swings between 6.91 and 44.42 token/s (52.7% CV).\n- No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage, so the four-hour window is complete.","summary_en":"- deepseek-v4.1-flash is the strongest model at 168.93 token/s average throughput, peaking at 230.47 token/s at 20:20 UTC; nemotron-3-ultra is the weakest at 24.86 token/s average, ending at 6.91 token/s at 23:00 UTC.\n- The most operationally significant volatility is deepseek-v4.1-flash collapsing to 15.72 token/s at 22:00 UTC from 215.07 token/s at 21:40, driving a 36.3% coefficient of variation; nemotron-3-ultra also swings between 6.91 and 44.42 token/s (52.7% CV).\n- No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage, so the four-hour window is complete.","summary_items":["deepseek-v4.1-flash is the strongest model at 168.93 token/s average throughput, peaking at 230.47 token/s at 20:20 UTC; nemotron-3-ultra is the weakest at 24.86 token/s average, ending at 6.91 token/s at 23:00 UTC.","The most operationally significant volatility is deepseek-v4.1-flash collapsing to 15.72 token/s at 22:00 UTC from 215.07 token/s at 21:40, driving a 36.3% coefficient of variation; nemotron-3-ultra also swings between 6.91 and 44.42 token/s (52.7% CV).","No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage, so the four-hour window is complete."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 168.93 \u8bcd\u5143/\u79d2\uff0c\u5728 UTC 20:20 \u8fbe\u5230\u5cf0\u503c 230.47 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 24.86 \u8bcd\u5143/\u79d2\uff0c\u5728 UTC 23:00 \u4ee5 6.91 \u8bcd\u5143/\u79d2\u6536\u5c3e\u3002\n- \u8fd0\u8425\u5c42\u9762\u6700\u663e\u8457\u7684\u6ce2\u52a8\u662f deepseek-v4.1-flash \u4ece 21:40 \u7684 215.07 \u8bcd\u5143/\u79d2\u9aa4\u964d\u81f3 22:00 \u7684 15.72 \u8bcd\u5143/\u79d2\uff0c\u5bfc\u81f4\u53d8\u5f02\u7cfb\u6570\u8fbe 36.3%\uff1bnemotron-3-ultra \u4e5f\u5728 6.91 \u81f3 44.42 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\uff08\u53d8\u5f02\u7cfb\u6570 52.7%\uff09\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62a5\u544a\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\u300196 \u4e2a\u6709\u6548\u6570\u636e\u70b9\u4ee5\u53ca 100.0% \u7684\u8986\u76d6\u7387\uff0c\u56e0\u6b64\u8be5\u56db\u5c0f\u65f6\u65f6\u95f4\u7a97\u53e3\u662f\u5b8c\u6574\u7684\u3002","total_duration_ns":2618616478,"translated_at":"2026-10-07T23:03:05.688098+00:00","translation_eval_count":230,"translation_model":"glm-5.3","translation_prompt_eval_count":365,"translation_total_duration_ns":927415004,"translation_wall_seconds":1.073,"valid_point_count":96,"wall_seconds":2.824},{"coverage_pct":100.0,"eval_count":350,"generated_at":"2026-10-07T22:03:07.753029+00:00","generated_label":"Oct 07, 2026 \u00b7 22:03 UTC","id":1205,"model":"glm-5.3","period_end":"2026-10-07T22:03:01+00:00","period_label":"Oct 07, 2026 \u00b7 18:03 UTC to Oct 07, 2026 \u00b7 22:03 UTC","period_start":"2026-10-07T18:03:01+00:00","prompt_eval_count":2246,"summary":"- deepseek-v4.1-flash is the strongest model at 174.19 token/s average throughput, while nemotron-3-ultra is the weakest at 23.52 token/s average; the next-lowest, minimax-m3, averaged 60.48 token/s.\n- The most significant event is deepseek-v4.1-flash collapsing from 215.07 token/s at 21:40 to 15.72 token/s at 22:00, versus a 230.47 token/s peak at 20:20. minimax-m3 also declined 26.4% over the window, and gemma4:31b swung between 33.24 and 164.52 token/s.\n- No missing-data limitation exists: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage for the four-hour period.","summary_en":"- deepseek-v4.1-flash is the strongest model at 174.19 token/s average throughput, while nemotron-3-ultra is the weakest at 23.52 token/s average; the next-lowest, minimax-m3, averaged 60.48 token/s.\n- The most significant event is deepseek-v4.1-flash collapsing from 215.07 token/s at 21:40 to 15.72 token/s at 22:00, versus a 230.47 token/s peak at 20:20. minimax-m3 also declined 26.4% over the window, and gemma4:31b swung between 33.24 and 164.52 token/s.\n- No missing-data limitation exists: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage for the four-hour period.","summary_items":["deepseek-v4.1-flash is the strongest model at 174.19 token/s average throughput, while nemotron-3-ultra is the weakest at 23.52 token/s average; the next-lowest, minimax-m3, averaged 60.48 token/s.","The most significant event is deepseek-v4.1-flash collapsing from 215.07 token/s at 21:40 to 15.72 token/s at 22:00, versus a 230.47 token/s peak at 20:20. minimax-m3 also declined 26.4% over the window, and gemma4:31b swung between 33.24 and 164.52 token/s.","No missing-data limitation exists: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage for the four-hour period."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 174.19 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 23.52 \u8bcd\u5143/\u79d2\uff1b\u6b21\u4f4e\u7684 minimax-m3 \u5e73\u5747\u503c\u4e3a 60.48 \u8bcd\u5143/\u79d2\u3002\n- \u6700\u663e\u8457\u7684\u4e8b\u4ef6\u662f deepseek-v4.1-flash \u4ece 21:40 \u7684 215.07 \u8bcd\u5143/\u79d2\u9aa4\u964d\u81f3 22:00 \u7684 15.72 \u8bcd\u5143/\u79d2\uff0c\u76f8\u6bd4\u4e4b\u4e0b\u5176\u5cf0\u503c\u51fa\u73b0\u5728 20:20\uff0c\u4e3a 230.47 \u8bcd\u5143/\u79d2\u3002minimax-m3 \u5728\u8be5\u65f6\u6bb5\u5185\u4e5f\u4e0b\u964d\u4e86 26.4%\uff0cgemma4:31b \u5219\u5728 33.24 \u81f3 164.52 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62a5\u544a\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\u300196 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u4e14\u5728\u56db\u5c0f\u65f6\u671f\u95f4\u5185\u8986\u76d6\u7387\u8fbe\u5230 100.0%\u3002","total_duration_ns":3322168532,"translated_at":"2026-10-07T22:03:07.753029+00:00","translation_eval_count":209,"translation_model":"glm-5.3","translation_prompt_eval_count":360,"translation_total_duration_ns":2261141423,"translation_wall_seconds":2.589,"valid_point_count":96,"wall_seconds":3.489},{"coverage_pct":100.0,"eval_count":469,"generated_at":"2026-10-07T21:03:08.957246+00:00","generated_label":"Oct 07, 2026 \u00b7 21:03 UTC","id":1204,"model":"glm-5.3","period_end":"2026-10-07T21:03:02+00:00","period_label":"Oct 07, 2026 \u00b7 17:03 UTC to Oct 07, 2026 \u00b7 21:03 UTC","period_start":"2026-10-07T17:03:02+00:00","prompt_eval_count":2247,"summary":"- deepseek-v4.1-flash is the strongest model at 198.05 token/s average throughput, while nemotron-3-ultra is the weakest at 25.11 token/s average, a gap of roughly eight times.\n- glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 72.5% and swings from 11.64 token/s at 17:20 to 195.82 token/s at 18:00; gemma4:31b likewise oscillates between 33.24 and 148.63 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 198.05 token/s average throughput, while nemotron-3-ultra is the weakest at 25.11 token/s average, a gap of roughly eight times.\n- glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 72.5% and swings from 11.64 token/s at 17:20 to 195.82 token/s at 18:00; gemma4:31b likewise oscillates between 33.24 and 148.63 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 198.05 token/s average throughput, while nemotron-3-ultra is the weakest at 25.11 token/s average, a gap of roughly eight times.","glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 72.5% and swings from 11.64 token/s at 17:20 to 195.82 token/s at 18:00; gemma4:31b likewise oscillates between 33.24 and 148.63 token/s.","No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6027\u80fd\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 198.05 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u662f\u6027\u80fd\u6700\u5f31\u7684\u6a21\u578b\uff0c\u5e73\u5747\u4e3a 25.11 \u8bcd\u5143/\u79d2\uff0c\u4e24\u8005\u5dee\u8ddd\u7ea6\u4e3a\u516b\u500d\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u5bf9\u8fd0\u884c\u5f71\u54cd\u6700\u663e\u8457\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 72.5%\uff0c\u541e\u5410\u91cf\u5728 17:20 \u7684 11.64 \u8bcd\u5143/\u79d2\u4e0e 18:00 \u7684 195.82 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6446\u52a8\uff1bgemma4:31b \u540c\u6837\u5728 33.24 \u81f3 148.63 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u632f\u8361\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u6570\u636e\u96c6\u62a5\u544a\u4e86 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u8986\u76d6\u7387\u8fbe\u5230 100.0%\u3002","total_duration_ns":4322557544,"translated_at":"2026-10-07T21:03:08.957246+00:00","translation_eval_count":188,"translation_model":"glm-5.3","translation_prompt_eval_count":335,"translation_total_duration_ns":2121205668,"translation_wall_seconds":2.275,"valid_point_count":96,"wall_seconds":4.655},{"coverage_pct":100.0,"eval_count":737,"generated_at":"2026-10-07T20:03:10.179554+00:00","generated_label":"Oct 07, 2026 \u00b7 20:03 UTC","id":1203,"model":"glm-5.3","period_end":"2026-10-07T20:03:01+00:00","period_label":"Oct 07, 2026 \u00b7 16:03 UTC to Oct 07, 2026 \u00b7 20:03 UTC","period_start":"2026-10-07T16:03:01+00:00","prompt_eval_count":2246,"summary":"- deepseek-v4.1-flash is the strongest model at 193.86 token/s average throughput, while nemotron-3-ultra is the weakest at 22.44 token/s average, roughly 8.6x lower.\n- glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 86.7% and throughput swinging from 11.64 token/s at 17:20 to 195.82 token/s at 18:00; glm-5.3 also declined 29.3% over the window.\n- No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 193.86 token/s average throughput, while nemotron-3-ultra is the weakest at 22.44 token/s average, roughly 8.6x lower.\n- glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 86.7% and throughput swinging from 11.64 token/s at 17:20 to 195.82 token/s at 18:00; glm-5.3 also declined 29.3% over the window.\n- No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 193.86 token/s average throughput, while nemotron-3-ultra is the weakest at 22.44 token/s average, roughly 8.6x lower.","glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 86.7% and throughput swinging from 11.64 token/s at 17:20 to 195.82 token/s at 18:00; glm-5.3 also declined 29.3% over the window.","No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 193.86 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u662f\u6700\u5f31\u6a21\u578b\uff0c\u5e73\u5747\u4e3a 22.44 \u8bcd\u5143/\u79d2\uff0c\u4f4e\u4e86\u7ea6 8.6 \u500d\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u5bf9\u8fd0\u8425\u5f71\u54cd\u6700\u663e\u8457\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 86.7%\uff0c\u541e\u5410\u91cf\u5728 17:20 \u7684 11.64 \u8bcd\u5143/\u79d2\u4e0e 18:00 \u7684 195.82 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u5927\u5e45\u6446\u52a8\uff1bglm-5.3 \u5728\u8be5\u65f6\u95f4\u7a97\u53e3\u5185\u4e5f\u4e0b\u964d\u4e86 29.3%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12/12 \u4e2a\u6837\u672c\uff0cvalid_point_count \u4e3a 96\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":6324883775,"translated_at":"2026-10-07T20:03:10.179554+00:00","translation_eval_count":178,"translation_model":"glm-5.3","translation_prompt_eval_count":332,"translation_total_duration_ns":2028061750,"translation_wall_seconds":2.177,"valid_point_count":96,"wall_seconds":6.488},{"coverage_pct":100.0,"eval_count":924,"generated_at":"2026-10-07T19:04:15.799467+00:00","generated_label":"Oct 07, 2026 \u00b7 19:04 UTC","id":1202,"model":"glm-5.3","period_end":"2026-10-07T19:03:01+00:00","period_label":"Oct 07, 2026 \u00b7 15:03 UTC to Oct 07, 2026 \u00b7 19:03 UTC","period_start":"2026-10-07T15:03:01+00:00","prompt_eval_count":2244,"summary":"- deepseek-v4.1-flash is the strongest model at 183.22 token/s average throughput (range 131.0\u2013220.41 token/s), while nemotron-3-ultra is the weakest at 18.41 token/s average (range 5.52\u201330.73 token/s).\n- glm-5.2 shows the most extreme volatility, with a 95.7% coefficient of variation, swinging from 11.64 token/s at 17:20 to 195.82 token/s at 18:00; glm-5.3 declined 32.5% over the window, ending at 57.77 token/s versus its 173.98 token/s peak.\n- No missing-data limitation applies: all eight models have 12 of 12 samples and 100.0% coverage, with 96 valid points, so results reflect the full four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 183.22 token/s average throughput (range 131.0\u2013220.41 token/s), while nemotron-3-ultra is the weakest at 18.41 token/s average (range 5.52\u201330.73 token/s).\n- glm-5.2 shows the most extreme volatility, with a 95.7% coefficient of variation, swinging from 11.64 token/s at 17:20 to 195.82 token/s at 18:00; glm-5.3 declined 32.5% over the window, ending at 57.77 token/s versus its 173.98 token/s peak.\n- No missing-data limitation applies: all eight models have 12 of 12 samples and 100.0% coverage, with 96 valid points, so results reflect the full four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 183.22 token/s average throughput (range 131.0\u2013220.41 token/s), while nemotron-3-ultra is the weakest at 18.41 token/s average (range 5.52\u201330.73 token/s).","glm-5.2 shows the most extreme volatility, with a 95.7% coefficient of variation, swinging from 11.64 token/s at 17:20 to 195.82 token/s at 18:00; glm-5.3 declined 32.5% over the window, ending at 57.77 token/s versus its 173.98 token/s peak.","No missing-data limitation applies: all eight models have 12 of 12 samples and 100.0% coverage, with 96 valid points, so results reflect the full four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 183.22 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 131.0\u2013220.41 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 18.41 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 5.52\u201330.73 \u8bcd\u5143/\u79d2\uff09\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u6700\u5267\u70c8\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u8fbe 95.7%\uff0c\u4ece 17:20 \u7684 11.64 \u8bcd\u5143/\u79d2\u6446\u52a8\u5230 18:00 \u7684 195.82 \u8bcd\u5143/\u79d2\uff1bglm-5.3 \u5728\u8be5\u65f6\u95f4\u7a97\u53e3\u5185\u4e0b\u964d\u4e86 32.5%\uff0c\u6700\u7ec8\u4e3a 57.77 \u8bcd\u5143/\u79d2\uff0c\u800c\u5176\u5cf0\u503c\u4e3a 173.98 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u56e0\u6b64\u7ed3\u679c\u53cd\u6620\u4e86\u5b8c\u6574\u7684\u56db\u5c0f\u65f6\u65f6\u95f4\u7a97\u53e3\u3002","total_duration_ns":72042718432,"translated_at":"2026-10-07T19:04:15.799467+00:00","translation_eval_count":218,"translation_model":"glm-5.3","translation_prompt_eval_count":361,"translation_total_duration_ns":1175172812,"translation_wall_seconds":1.493,"valid_point_count":96,"wall_seconds":72.396},{"coverage_pct":100.0,"eval_count":322,"generated_at":"2026-10-07T18:03:09.998016+00:00","generated_label":"Oct 07, 2026 \u00b7 18:03 UTC","id":1201,"model":"glm-5.3","period_end":"2026-10-07T18:03:01+00:00","period_label":"Oct 07, 2026 \u00b7 14:03 UTC to Oct 07, 2026 \u00b7 18:03 UTC","period_start":"2026-10-07T14:03:01+00:00","prompt_eval_count":2245,"summary":"- deepseek-v4.1-flash is the strongest model at 184.03 token/s average throughput (range 131.0 to 228.29 token/s), while nemotron-3-ultra is the weakest at 19.65 token/s average (range 5.52 to 30.73 token/s).\n- The most significant volatility is glm-5.2, whose coefficient of variation is 100.0%: it held roughly 20 to 46 token/s for eleven intervals, then spiked to 195.82 token/s at 18:00, far above its 46.06 token/s average. glm-5.3-flash also shows a strong upward trend of 82.4%, rising from 69.94 to 154.04 token/s.\n- No missing-data limitation exists: all eight models have 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 184.03 token/s average throughput (range 131.0 to 228.29 token/s), while nemotron-3-ultra is the weakest at 19.65 token/s average (range 5.52 to 30.73 token/s).\n- The most significant volatility is glm-5.2, whose coefficient of variation is 100.0%: it held roughly 20 to 46 token/s for eleven intervals, then spiked to 195.82 token/s at 18:00, far above its 46.06 token/s average. glm-5.3-flash also shows a strong upward trend of 82.4%, rising from 69.94 to 154.04 token/s.\n- No missing-data limitation exists: all eight models have 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 184.03 token/s average throughput (range 131.0 to 228.29 token/s), while nemotron-3-ultra is the weakest at 19.65 token/s average (range 5.52 to 30.73 token/s).","The most significant volatility is glm-5.2, whose coefficient of variation is 100.0%: it held roughly 20 to 46 token/s for eleven intervals, then spiked to 195.82 token/s at 18:00, far above its 46.06 token/s average. glm-5.3-flash also shows a strong upward trend of 82.4%, rising from 69.94 to 154.04 token/s.","No missing-data limitation exists: all eight models have 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 184.03 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 131.0 \u81f3 228.29 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 19.65 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 5.52 \u81f3 30.73 \u8bcd\u5143/\u79d2\uff09\u3002\n- \u6ce2\u52a8\u6700\u663e\u8457\u7684\u662f glm-5.2\uff0c\u5176\u53d8\u5f02\u7cfb\u6570\u4e3a 100.0%\uff1a\u5b83\u5728\u5341\u4e00\u4e2a\u65f6\u95f4\u533a\u95f4\u5185\u4fdd\u6301\u5728\u7ea6 20 \u81f3 46 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u5728 18:00 \u98d9\u5347\u81f3 195.82 \u8bcd\u5143/\u79d2\uff0c\u8fdc\u9ad8\u4e8e\u5176 46.06 \u8bcd\u5143/\u79d2\u7684\u5e73\u5747\u503c\u3002glm-5.3-flash \u4e5f\u5448\u73b0\u51fa 82.4% \u7684\u5f3a\u52b2\u4e0a\u5347\u8d8b\u52bf\uff0c\u4ece 69.94 \u5347\u81f3 154.04 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62e5\u6709 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\u300196 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5e76\u4e14\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u8fbe\u5230 100.0% \u7684\u8986\u76d6\u7387\u3002","total_duration_ns":5950004203,"translated_at":"2026-10-07T18:03:09.998016+00:00","translation_eval_count":231,"translation_model":"glm-5.3","translation_prompt_eval_count":373,"translation_total_duration_ns":2163930857,"translation_wall_seconds":2.308,"valid_point_count":96,"wall_seconds":6.117},{"coverage_pct":100.0,"eval_count":758,"generated_at":"2026-10-07T17:03:07.400770+00:00","generated_label":"Oct 07, 2026 \u00b7 17:03 UTC","id":1200,"model":"glm-5.3","period_end":"2026-10-07T17:03:02+00:00","period_label":"Oct 07, 2026 \u00b7 13:03 UTC to Oct 07, 2026 \u00b7 17:03 UTC","period_start":"2026-10-07T13:03:02+00:00","prompt_eval_count":2246,"summary":"- deepseek-v4.1-flash is the strongest model at 185.14 token/s average throughput (range 131.0 to 229.6 token/s), while nemotron-3-ultra is the weakest at 18.76 token/s average (range 9.79 to 30.73 token/s).\n- gemma4:31b shows the highest volatility, with a 42.7% coefficient of variation and swings between 27.01 and 159.28 token/s; glm-5.2 shows the steepest decline, trending down 30.2% to a 32.82 token/s latest reading.\n- No missing-data limitation: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 185.14 token/s average throughput (range 131.0 to 229.6 token/s), while nemotron-3-ultra is the weakest at 18.76 token/s average (range 9.79 to 30.73 token/s).\n- gemma4:31b shows the highest volatility, with a 42.7% coefficient of variation and swings between 27.01 and 159.28 token/s; glm-5.2 shows the steepest decline, trending down 30.2% to a 32.82 token/s latest reading.\n- No missing-data limitation: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 185.14 token/s average throughput (range 131.0 to 229.6 token/s), while nemotron-3-ultra is the weakest at 18.76 token/s average (range 9.79 to 30.73 token/s).","gemma4:31b shows the highest volatility, with a 42.7% coefficient of variation and swings between 27.01 and 159.28 token/s; glm-5.2 shows the steepest decline, trending down 30.2% to a 32.82 token/s latest reading.","No missing-data limitation: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 185.14 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 131.0 \u81f3 229.6 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 18.76 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 9.79 \u81f3 30.73 \u8bcd\u5143/\u79d2\uff09\u3002\n- gemma4:31b \u6ce2\u52a8\u6027\u6700\u9ad8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 42.7%\uff0c\u5728 27.01 \u81f3 159.28 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\uff1bglm-5.2 \u4e0b\u964d\u5e45\u5ea6\u6700\u9661\uff0c\u5448 30.2% \u7684\u4e0b\u964d\u8d8b\u52bf\uff0c\u6700\u65b0\u8bfb\u6570\u4e3a 32.82 \u8bcd\u5143/\u79d2\u3002\n- \u65e0\u7f3a\u5931\u6570\u636e\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u8986\u76d6\u7387\u8fbe\u5230 100.0%\u3002","total_duration_ns":3906296901,"translated_at":"2026-10-07T17:03:07.400770+00:00","translation_eval_count":192,"translation_model":"glm-5.3","translation_prompt_eval_count":344,"translation_total_duration_ns":919713252,"translation_wall_seconds":1.241,"valid_point_count":96,"wall_seconds":4.062},{"coverage_pct":100.0,"eval_count":478,"generated_at":"2026-10-07T16:03:05.452279+00:00","generated_label":"Oct 07, 2026 \u00b7 16:03 UTC","id":1199,"model":"glm-5.3","period_end":"2026-10-07T16:03:01+00:00","period_label":"Oct 07, 2026 \u00b7 12:03 UTC to Oct 07, 2026 \u00b7 16:03 UTC","period_start":"2026-10-07T12:03:01+00:00","prompt_eval_count":2247,"summary":"- deepseek-v4.1-flash is the strongest model at 187.92 token/s average throughput (peak 229.6 token/s), while nemotron-3-ultra is the weakest at 16.56 token/s average (peak 25.16 token/s).\n- gemma4:31b shows the highest volatility, with a coefficient of variation of 56.6% and swings between 27.01 and 159.28 token/s within the four hours; glm-5.2 declined 51.0% from 99.08 to 21.99 token/s.\n- No missing-data limitation exists: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, so the averages are based on complete data.","summary_en":"- deepseek-v4.1-flash is the strongest model at 187.92 token/s average throughput (peak 229.6 token/s), while nemotron-3-ultra is the weakest at 16.56 token/s average (peak 25.16 token/s).\n- gemma4:31b shows the highest volatility, with a coefficient of variation of 56.6% and swings between 27.01 and 159.28 token/s within the four hours; glm-5.2 declined 51.0% from 99.08 to 21.99 token/s.\n- No missing-data limitation exists: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, so the averages are based on complete data.","summary_items":["deepseek-v4.1-flash is the strongest model at 187.92 token/s average throughput (peak 229.6 token/s), while nemotron-3-ultra is the weakest at 16.56 token/s average (peak 25.16 token/s).","gemma4:31b shows the highest volatility, with a coefficient of variation of 56.6% and swings between 27.01 and 159.28 token/s within the four hours; glm-5.2 declined 51.0% from 99.08 to 21.99 token/s.","No missing-data limitation exists: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, so the averages are based on complete data."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 187.92 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 229.6 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 16.56 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 25.16 \u8bcd\u5143/\u79d2\uff09\u3002\n- gemma4:31b \u7684\u6ce2\u52a8\u6027\u6700\u9ad8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 56.6%\uff0c\u5728\u56db\u5c0f\u65f6\u5185\u4e8e 27.01 \u81f3 159.28 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\uff1bglm-5.2 \u4ece 99.08 \u4e0b\u964d\u81f3 21.99 \u8bcd\u5143/\u79d2\uff0c\u964d\u5e45\u4e3a 51.0%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u56e0\u6b64\u5e73\u5747\u503c\u57fa\u4e8e\u5b8c\u6574\u6570\u636e\u8ba1\u7b97\u3002","total_duration_ns":2689343483,"translated_at":"2026-10-07T16:03:05.452279+00:00","translation_eval_count":196,"translation_model":"glm-5.3","translation_prompt_eval_count":338,"translation_total_duration_ns":1216074889,"translation_wall_seconds":1.363,"valid_point_count":96,"wall_seconds":2.859},{"coverage_pct":100.0,"eval_count":878,"generated_at":"2026-10-07T15:03:14.467683+00:00","generated_label":"Oct 07, 2026 \u00b7 15:03 UTC","id":1198,"model":"glm-5.3","period_end":"2026-10-07T15:03:02+00:00","period_label":"Oct 07, 2026 \u00b7 11:03 UTC to Oct 07, 2026 \u00b7 15:03 UTC","period_start":"2026-10-07T11:03:02+00:00","prompt_eval_count":2249,"summary":"- deepseek-v4.1-flash is the strongest model at 194.19 token/s average throughput (range 141.51\u2013229.60 token/s), while nemotron-3-ultra is the weakest at 19.0 token/s average (range 7.34\u201339.55 token/s).\n- The most operationally significant movement is glm-5.2's steady decline, with a -29.3% trend from a 102.89 token/s peak at 12:00 UTC to 31.99 token/s at 15:00; gemma4:31b shows the highest volatility with a 50.5% coefficient of variation, swinging between 30.28 and 146.33 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 194.19 token/s average throughput (range 141.51\u2013229.60 token/s), while nemotron-3-ultra is the weakest at 19.0 token/s average (range 7.34\u201339.55 token/s).\n- The most operationally significant movement is glm-5.2's steady decline, with a -29.3% trend from a 102.89 token/s peak at 12:00 UTC to 31.99 token/s at 15:00; gemma4:31b shows the highest volatility with a 50.5% coefficient of variation, swinging between 30.28 and 146.33 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 194.19 token/s average throughput (range 141.51\u2013229.60 token/s), while nemotron-3-ultra is the weakest at 19.0 token/s average (range 7.34\u201339.55 token/s).","The most operationally significant movement is glm-5.2's steady decline, with a -29.3% trend from a 102.89 token/s peak at 12:00 UTC to 31.99 token/s at 15:00; gemma4:31b shows the highest volatility with a 50.5% coefficient of variation, swinging between 30.28 and 146.33 token/s.","No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 194.19 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 141.51\u2013229.60 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u662f\u6700\u5f31\u6a21\u578b\uff0c\u5e73\u5747\u4e3a 19.0 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 7.34\u201339.55 \u8bcd\u5143/\u79d2\uff09\u3002\n- \u8fd0\u8425\u5c42\u9762\u6700\u663e\u8457\u7684\u53d8\u5316\u662f glm-5.2 \u7684\u6301\u7eed\u4e0b\u6ed1\uff0c\u8d8b\u52bf\u4e3a -29.3%\uff0c\u4ece 12:00 UTC \u7684 102.89 \u8bcd\u5143/\u79d2\u5cf0\u503c\u964d\u81f3 15:00 \u7684 31.99 \u8bcd\u5143/\u79d2\uff1bgemma4:31b \u6ce2\u52a8\u6027\u6700\u9ad8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 50.5%\uff0c\u5728 30.28 \u548c 146.33 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12/12 \u4e2a\u6837\u672c\uff0cvalid_point_count \u4e3a 96\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":5748928180,"translated_at":"2026-10-07T15:03:14.467683+00:00","translation_eval_count":212,"translation_model":"glm-5.3","translation_prompt_eval_count":368,"translation_total_duration_ns":6344157062,"translation_wall_seconds":6.496,"valid_point_count":96,"wall_seconds":5.908},{"coverage_pct":100.0,"eval_count":745,"generated_at":"2026-10-07T14:03:07.516341+00:00","generated_label":"Oct 07, 2026 \u00b7 14:03 UTC","id":1197,"model":"glm-5.3","period_end":"2026-10-07T14:03:02+00:00","period_label":"Oct 07, 2026 \u00b7 10:03 UTC to Oct 07, 2026 \u00b7 14:03 UTC","period_start":"2026-10-07T10:03:02+00:00","prompt_eval_count":2253,"summary":"- deepseek-v4.1-flash is the strongest model at 199.16 token/s average throughput (range 166.81\u2013229.76 token/s), while nemotron-3-ultra is the weakest at 19.77 token/s average (range 7.34\u201339.55 token/s).\n- gemma4:31b shows the most operationally significant volatility, with a 48.1% coefficient of variation, swings between 153.24 and 30.28 token/s within the window, and a -41.9% trend; nemotron-3-ultra also declined 45.8%.\n- No missing-data limitation applies: all eight models report 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0%, so every twenty-minute observation across the four-hour window is present.","summary_en":"- deepseek-v4.1-flash is the strongest model at 199.16 token/s average throughput (range 166.81\u2013229.76 token/s), while nemotron-3-ultra is the weakest at 19.77 token/s average (range 7.34\u201339.55 token/s).\n- gemma4:31b shows the most operationally significant volatility, with a 48.1% coefficient of variation, swings between 153.24 and 30.28 token/s within the window, and a -41.9% trend; nemotron-3-ultra also declined 45.8%.\n- No missing-data limitation applies: all eight models report 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0%, so every twenty-minute observation across the four-hour window is present.","summary_items":["deepseek-v4.1-flash is the strongest model at 199.16 token/s average throughput (range 166.81\u2013229.76 token/s), while nemotron-3-ultra is the weakest at 19.77 token/s average (range 7.34\u201339.55 token/s).","gemma4:31b shows the most operationally significant volatility, with a 48.1% coefficient of variation, swings between 153.24 and 30.28 token/s within the window, and a -41.9% trend; nemotron-3-ultra also declined 45.8%.","No missing-data limitation applies: all eight models report 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0%, so every twenty-minute observation across the four-hour window is present."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 199.16 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 166.81\u2013229.76 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 19.77 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 7.34\u201339.55 \u8bcd\u5143/\u79d2\uff09\u3002\n- gemma4:31b \u8868\u73b0\u51fa\u6700\u5177\u8fd0\u8425\u610f\u4e49\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 48.1%\uff0c\u5728\u7a97\u53e3\u5185\u4e8e 153.24 \u548c 30.28 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6446\u52a8\uff0c\u8d8b\u52bf\u4e3a -41.9%\uff1bnemotron-3-ultra \u4e5f\u4e0b\u964d\u4e86 45.8%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62a5\u544a\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0cvalid_point_count \u4e3a 96\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u56e0\u6b64\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u7684\u6bcf\u4e8c\u5341\u5206\u949f\u89c2\u6d4b\u503c\u5747\u5b8c\u6574\u5b58\u5728\u3002","total_duration_ns":4033009542,"translated_at":"2026-10-07T14:03:07.516341+00:00","translation_eval_count":213,"translation_model":"glm-5.3","translation_prompt_eval_count":353,"translation_total_duration_ns":1160913670,"translation_wall_seconds":1.304,"valid_point_count":96,"wall_seconds":4.205},{"coverage_pct":100.0,"eval_count":995,"generated_at":"2026-10-07T13:03:08.782927+00:00","generated_label":"Oct 07, 2026 \u00b7 13:03 UTC","id":1196,"model":"glm-5.3","period_end":"2026-10-07T13:03:01+00:00","period_label":"Oct 07, 2026 \u00b7 09:03 UTC to Oct 07, 2026 \u00b7 13:03 UTC","period_start":"2026-10-07T09:03:01+00:00","prompt_eval_count":2250,"summary":"- deepseek-v4.1-flash is the strongest model at 195.11 token/s average throughput (range 166.81\u2013229.76 token/s), while nemotron-3-ultra is the weakest at 18.84 token/s average (range 7.34\u201339.55 token/s).\n- The most operationally significant movement is gemma4:31b's decline of 38.3 percent, falling from 153.24 token/s at 11:00 to 30.28 token/s at 13:00; minimax-m3 also dropped to 6.99 token/s at 12:00, and nemotron-3-ultra shows the highest volatility at 59.0 percent coefficient of variation.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0 percent, so the four-hour window is complete.","summary_en":"- deepseek-v4.1-flash is the strongest model at 195.11 token/s average throughput (range 166.81\u2013229.76 token/s), while nemotron-3-ultra is the weakest at 18.84 token/s average (range 7.34\u201339.55 token/s).\n- The most operationally significant movement is gemma4:31b's decline of 38.3 percent, falling from 153.24 token/s at 11:00 to 30.28 token/s at 13:00; minimax-m3 also dropped to 6.99 token/s at 12:00, and nemotron-3-ultra shows the highest volatility at 59.0 percent coefficient of variation.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0 percent, so the four-hour window is complete.","summary_items":["deepseek-v4.1-flash is the strongest model at 195.11 token/s average throughput (range 166.81\u2013229.76 token/s), while nemotron-3-ultra is the weakest at 18.84 token/s average (range 7.34\u201339.55 token/s).","The most operationally significant movement is gemma4:31b's decline of 38.3 percent, falling from 153.24 token/s at 11:00 to 30.28 token/s at 13:00; minimax-m3 also dropped to 6.99 token/s at 12:00, and nemotron-3-ultra shows the highest volatility at 59.0 percent coefficient of variation.","No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0 percent, so the four-hour window is complete."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 195.11 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 166.81\u2013229.76 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 18.84 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 7.34\u201339.55 \u8bcd\u5143/\u79d2\uff09\u3002\n- \u8fd0\u8425\u5c42\u9762\u6700\u663e\u8457\u7684\u53d8\u5316\u662f gemma4:31b \u4e0b\u964d\u4e86 38.3%\uff0c\u4ece 11:00 \u7684 153.24 \u8bcd\u5143/\u79d2\u964d\u81f3 13:00 \u7684 30.28 \u8bcd\u5143/\u79d2\uff1bminimax-m3 \u5728 12:00 \u4e5f\u964d\u81f3 6.99 \u8bcd\u5143/\u79d2\uff0c\u4e14 nemotron-3-ultra \u7684\u6ce2\u52a8\u6027\u6700\u9ad8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 59.0%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0cvalid_point_count \u4e3a 96\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u56e0\u6b64\u56db\u5c0f\u65f6\u7a97\u53e3\u662f\u5b8c\u6574\u7684\u3002","total_duration_ns":4735626706,"translated_at":"2026-10-07T13:03:08.782927+00:00","translation_eval_count":247,"translation_model":"glm-5.3","translation_prompt_eval_count":371,"translation_total_duration_ns":1958838440,"translation_wall_seconds":2.103,"valid_point_count":96,"wall_seconds":5.069},{"coverage_pct":100.0,"eval_count":917,"generated_at":"2026-10-07T12:03:07.643034+00:00","generated_label":"Oct 07, 2026 \u00b7 12:03 UTC","id":1195,"model":"glm-5.3","period_end":"2026-10-07T12:03:01+00:00","period_label":"Oct 07, 2026 \u00b7 08:03 UTC to Oct 07, 2026 \u00b7 12:03 UTC","period_start":"2026-10-07T08:03:01+00:00","prompt_eval_count":2250,"summary":"- deepseek-v4.1-flash is the strongest model at 190.08 token/s average throughput (range 161.96\u2013229.76 token/s), while nemotron-3-ultra is the weakest at 19.9 token/s average (range 9.3\u201339.55 token/s).\n- The sharpest operational swing is minimax-m3's collapse from 85.57 token/s at 11:40 to 6.99 token/s at 12:00, its minimum; nemotron-3-ultra similarly fell from 39.55 to 10.36 token/s in the same interval. glm-5.2 shows the highest relative volatility (cv 57.6%), swinging between 14.66 and 134.7 token/s.\n- No missing-data limitation exists: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage over the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 190.08 token/s average throughput (range 161.96\u2013229.76 token/s), while nemotron-3-ultra is the weakest at 19.9 token/s average (range 9.3\u201339.55 token/s).\n- The sharpest operational swing is minimax-m3's collapse from 85.57 token/s at 11:40 to 6.99 token/s at 12:00, its minimum; nemotron-3-ultra similarly fell from 39.55 to 10.36 token/s in the same interval. glm-5.2 shows the highest relative volatility (cv 57.6%), swinging between 14.66 and 134.7 token/s.\n- No missing-data limitation exists: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage over the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 190.08 token/s average throughput (range 161.96\u2013229.76 token/s), while nemotron-3-ultra is the weakest at 19.9 token/s average (range 9.3\u201339.55 token/s).","The sharpest operational swing is minimax-m3's collapse from 85.57 token/s at 11:40 to 6.99 token/s at 12:00, its minimum; nemotron-3-ultra similarly fell from 39.55 to 10.36 token/s in the same interval. glm-5.2 shows the highest relative volatility (cv 57.6%), swinging between 14.66 and 134.7 token/s.","No missing-data limitation exists: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage over the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 190.08 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 161.96\u2013229.76 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 19.9 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 9.3\u201339.55 \u8bcd\u5143/\u79d2\uff09\u3002\n- \u6700\u5267\u70c8\u7684\u8fd0\u884c\u6ce2\u52a8\u662f minimax-m3 \u4ece 11:40 \u7684 85.57 \u8bcd\u5143/\u79d2\u9aa4\u964d\u81f3 12:00 \u7684 6.99 \u8bcd\u5143/\u79d2\uff08\u5176\u6700\u4f4e\u503c\uff09\uff1bnemotron-3-ultra \u5728\u540c\u4e00\u65f6\u95f4\u6bb5\u5185\u540c\u6837\u4ece 39.55 \u964d\u81f3 10.36 \u8bcd\u5143/\u79d2\u3002glm-5.2 \u663e\u793a\u51fa\u6700\u9ad8\u7684\u76f8\u5bf9\u6ce2\u52a8\u6027\uff08\u53d8\u5f02\u7cfb\u6570 57.6%\uff09\uff0c\u5728 14.66 \u548c 134.7 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6446\u52a8\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u6570\u636e\u96c6\u62a5\u544a\u4e86 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u8986\u76d6\u7387\u8fbe\u5230 100.0%\u3002","total_duration_ns":4121844080,"translated_at":"2026-10-07T12:03:07.643034+00:00","translation_eval_count":234,"translation_model":"glm-5.3","translation_prompt_eval_count":376,"translation_total_duration_ns":1152675375,"translation_wall_seconds":1.297,"valid_point_count":96,"wall_seconds":4.471},{"coverage_pct":100.0,"eval_count":314,"generated_at":"2026-10-07T11:03:04.693597+00:00","generated_label":"Oct 07, 2026 \u00b7 11:03 UTC","id":1194,"model":"glm-5.3","period_end":"2026-10-07T11:03:01+00:00","period_label":"Oct 07, 2026 \u00b7 07:03 UTC to Oct 07, 2026 \u00b7 11:03 UTC","period_start":"2026-10-07T07:03:01+00:00","prompt_eval_count":2249,"summary":"- deepseek-v4.1-flash is the strongest model at 185.52 token/s average throughput (155.89\u2013229.76 token/s range), while nemotron-3-ultra is the weakest at 19.17 token/s average (9.3\u201339.15 token/s range).\n- glm-5.2 shows the most operationally significant volatility, with a 55.2% coefficient of variation and a -33.2% trend, falling from a 134.7 token/s peak at 09:00 to 14.66 token/s at 10:20; minimax-m3 also swung between 10.24 and 63.24 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, so the four-hour window is complete.","summary_en":"- deepseek-v4.1-flash is the strongest model at 185.52 token/s average throughput (155.89\u2013229.76 token/s range), while nemotron-3-ultra is the weakest at 19.17 token/s average (9.3\u201339.15 token/s range).\n- glm-5.2 shows the most operationally significant volatility, with a 55.2% coefficient of variation and a -33.2% trend, falling from a 134.7 token/s peak at 09:00 to 14.66 token/s at 10:20; minimax-m3 also swung between 10.24 and 63.24 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, so the four-hour window is complete.","summary_items":["deepseek-v4.1-flash is the strongest model at 185.52 token/s average throughput (155.89\u2013229.76 token/s range), while nemotron-3-ultra is the weakest at 19.17 token/s average (9.3\u201339.15 token/s range).","glm-5.2 shows the most operationally significant volatility, with a 55.2% coefficient of variation and a -33.2% trend, falling from a 134.7 token/s peak at 09:00 to 14.66 token/s at 10:20; minimax-m3 also swung between 10.24 and 63.24 token/s.","No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, so the four-hour window is complete."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 185.52 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 155.89\u2013229.76 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u662f\u6700\u5f31\u7684\u6a21\u578b\uff0c\u5e73\u5747\u4e3a 19.17 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 9.3\u201339.15 \u8bcd\u5143/\u79d2\uff09\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u5bf9\u8fd0\u8425\u5f71\u54cd\u6700\u5927\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 55.2%\uff0c\u8d8b\u52bf\u4e3a -33.2%\uff0c\u4ece 09:00 \u7684 134.7 \u8bcd\u5143/\u79d2\u5cf0\u503c\u964d\u81f3 10:20 \u7684 14.66 \u8bcd\u5143/\u79d2\uff1bminimax-m3 \u4e5f\u5728 10.24 \u81f3 63.24 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u56e0\u6b64\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u662f\u5b8c\u6574\u7684\u3002","total_duration_ns":1596707878,"translated_at":"2026-10-07T11:03:04.693597+00:00","translation_eval_count":228,"translation_model":"glm-5.3","translation_prompt_eval_count":357,"translation_total_duration_ns":1170597555,"translation_wall_seconds":1.507,"valid_point_count":96,"wall_seconds":1.766},{"coverage_pct":100.0,"eval_count":334,"generated_at":"2026-10-07T10:03:29.403432+00:00","generated_label":"Oct 07, 2026 \u00b7 10:03 UTC","id":1193,"model":"glm-5.3","period_end":"2026-10-07T10:03:01+00:00","period_label":"Oct 07, 2026 \u00b7 06:03 UTC to Oct 07, 2026 \u00b7 10:03 UTC","period_start":"2026-10-07T06:03:01+00:00","prompt_eval_count":2246,"summary":"- deepseek-v4.1-flash is the strongest model at 166.36 token/s average throughput (peak 225.27 token/s), while nemotron-3-ultra is the weakest at 18.17 token/s average (peak 28.23 token/s).\n- deepseek-v4.1-flash shows the strongest upward trend at +21.6%, versus nemotron-3-ultra at -36.2% and glm-5.2 at -31.4%; glm-5.2 is also the most volatile, with a 52.3% coefficient of variation spanning 31.79 to 159.68 token/s.\n- No missing-data limitation applies: all 8 models have 12 of 12 samples, 96 valid points, and 100.0% coverage, so the four-hour window is fully represented.","summary_en":"- deepseek-v4.1-flash is the strongest model at 166.36 token/s average throughput (peak 225.27 token/s), while nemotron-3-ultra is the weakest at 18.17 token/s average (peak 28.23 token/s).\n- deepseek-v4.1-flash shows the strongest upward trend at +21.6%, versus nemotron-3-ultra at -36.2% and glm-5.2 at -31.4%; glm-5.2 is also the most volatile, with a 52.3% coefficient of variation spanning 31.79 to 159.68 token/s.\n- No missing-data limitation applies: all 8 models have 12 of 12 samples, 96 valid points, and 100.0% coverage, so the four-hour window is fully represented.","summary_items":["deepseek-v4.1-flash is the strongest model at 166.36 token/s average throughput (peak 225.27 token/s), while nemotron-3-ultra is the weakest at 18.17 token/s average (peak 28.23 token/s).","deepseek-v4.1-flash shows the strongest upward trend at +21.6%, versus nemotron-3-ultra at -36.2% and glm-5.2 at -31.4%; glm-5.2 is also the most volatile, with a 52.3% coefficient of variation spanning 31.79 to 159.68 token/s.","No missing-data limitation applies: all 8 models have 12 of 12 samples, 96 valid points, and 100.0% coverage, so the four-hour window is fully represented."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 166.36 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 225.27 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 18.17 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 28.23 \u8bcd\u5143/\u79d2\uff09\u3002\n- deepseek-v4.1-flash \u5448\u73b0\u6700\u5f3a\u7684\u4e0a\u5347\u8d8b\u52bf\uff0c\u4e3a +21.6%\uff0c\u76f8\u6bd4\u4e4b\u4e0b nemotron-3-ultra \u4e3a -36.2%\uff0cglm-5.2 \u4e3a -31.4%\uff1bglm-5.2 \u540c\u65f6\u4e5f\u662f\u6ce2\u52a8\u6027\u6700\u5927\u7684\u6a21\u578b\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 52.3%\uff0c\u6ce2\u52a8\u8303\u56f4\u4ecb\u4e8e 31.79 \u81f3 159.68 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8 8 \u4e2a\u6a21\u578b\u5747\u6709 12/12 \u4e2a\u6837\u672c\u300196 \u4e2a\u6709\u6548\u6570\u636e\u70b9\u548c 100.0% \u7684\u8986\u76d6\u7387\uff0c\u56e0\u6b64\u8be5\u56db\u5c0f\u65f6\u7a97\u53e3\u5f97\u5230\u4e86\u5b8c\u6574\u5448\u73b0\u3002","total_duration_ns":7327333886,"translated_at":"2026-10-07T10:03:29.403432+00:00","translation_eval_count":196,"translation_model":"glm-5.3","translation_prompt_eval_count":350,"translation_total_duration_ns":20249653298,"translation_wall_seconds":20.396,"valid_point_count":96,"wall_seconds":7.491},{"coverage_pct":100.0,"eval_count":797,"generated_at":"2026-10-07T09:03:10.828570+00:00","generated_label":"Oct 07, 2026 \u00b7 09:03 UTC","id":1192,"model":"glm-5.3","period_end":"2026-10-07T09:03:01+00:00","period_label":"Oct 07, 2026 \u00b7 05:03 UTC to Oct 07, 2026 \u00b7 09:03 UTC","period_start":"2026-10-07T05:03:01+00:00","prompt_eval_count":2246,"summary":"- deepseek-v4.1-flash is the strongest model at 157.62 token/s average throughput (peak 224.58 token/s), while nemotron-3-ultra is the weakest at 21.43 token/s average, peaking at only 40.16 token/s.\n- glm-5.3-flash shows the most significant upward trend at +46.6%, climbing from 59.62 token/s at 06:40 to 164.11 token/s at 08:40, whereas minimax-m3 declined 39.2% to 10.24 token/s at 09:00; glm-5.2 is the most volatile with a 48.9% coefficient of variation.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage.","summary_en":"- deepseek-v4.1-flash is the strongest model at 157.62 token/s average throughput (peak 224.58 token/s), while nemotron-3-ultra is the weakest at 21.43 token/s average, peaking at only 40.16 token/s.\n- glm-5.3-flash shows the most significant upward trend at +46.6%, climbing from 59.62 token/s at 06:40 to 164.11 token/s at 08:40, whereas minimax-m3 declined 39.2% to 10.24 token/s at 09:00; glm-5.2 is the most volatile with a 48.9% coefficient of variation.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage.","summary_items":["deepseek-v4.1-flash is the strongest model at 157.62 token/s average throughput (peak 224.58 token/s), while nemotron-3-ultra is the weakest at 21.43 token/s average, peaking at only 40.16 token/s.","glm-5.3-flash shows the most significant upward trend at +46.6%, climbing from 59.62 token/s at 06:40 to 164.11 token/s at 08:40, whereas minimax-m3 declined 39.2% to 10.24 token/s at 09:00; glm-5.2 is the most volatile with a 48.9% coefficient of variation.","No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 157.62 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 224.58 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 21.43 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u4ec5\u4e3a 40.16 \u8bcd\u5143/\u79d2\u3002\n- glm-5.3-flash \u5448\u73b0\u6700\u663e\u8457\u7684\u4e0a\u5347\u8d8b\u52bf\uff0c\u8fbe +46.6%\uff0c\u4ece 06:40 \u7684 59.62 \u8bcd\u5143/\u79d2\u6500\u5347\u81f3 08:40 \u7684 164.11 \u8bcd\u5143/\u79d2\uff1b\u800c minimax-m3 \u4e0b\u964d\u4e86 39.2%\uff0c\u81f3 09:00 \u7684 10.24 \u8bcd\u5143/\u79d2\uff1bglm-5.2 \u6ce2\u52a8\u6027\u6700\u5927\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 48.9%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u6570\u636e\u96c6\u62a5\u544a\u4e86 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":7455200304,"translated_at":"2026-10-07T09:03:10.828570+00:00","translation_eval_count":219,"translation_model":"glm-5.3","translation_prompt_eval_count":358,"translation_total_duration_ns":1308440296,"translation_wall_seconds":1.466,"valid_point_count":96,"wall_seconds":7.622},{"coverage_pct":100.0,"eval_count":516,"generated_at":"2026-10-07T07:03:08.534556+00:00","generated_label":"Oct 07, 2026 \u00b7 07:03 UTC","id":1191,"model":"glm-5.3","period_end":"2026-10-07T07:03:01+00:00","period_label":"Oct 07, 2026 \u00b7 03:03 UTC to Oct 07, 2026 \u00b7 07:03 UTC","period_start":"2026-10-07T03:03:01+00:00","prompt_eval_count":2251,"summary":"- deepseek-v4.1-flash is the strongest model at 170.43 token/s average throughput, while nemotron-3-ultra is the weakest at 21.71 token/s, an eightfold gap between the two.\n- glm-5.2 shows the highest volatility (cv 69.8%), swinging between 14.83 and 195.72 token/s, including a low of 14.83 token/s at 06:00; deepseek-v4.1-flash also fell from 224.58 token/s at 06:40 to 35.29 token/s at 06:20, then recovered.\n- No missing-data limitation applies: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 170.43 token/s average throughput, while nemotron-3-ultra is the weakest at 21.71 token/s, an eightfold gap between the two.\n- glm-5.2 shows the highest volatility (cv 69.8%), swinging between 14.83 and 195.72 token/s, including a low of 14.83 token/s at 06:00; deepseek-v4.1-flash also fell from 224.58 token/s at 06:40 to 35.29 token/s at 06:20, then recovered.\n- No missing-data limitation applies: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 170.43 token/s average throughput, while nemotron-3-ultra is the weakest at 21.71 token/s, an eightfold gap between the two.","glm-5.2 shows the highest volatility (cv 69.8%), swinging between 14.83 and 195.72 token/s, including a low of 14.83 token/s at 06:00; deepseek-v4.1-flash also fell from 224.58 token/s at 06:40 to 35.29 token/s at 06:20, then recovered.","No missing-data limitation applies: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 170.43 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 21.71 \u8bcd\u5143/\u79d2\uff0c\u4e24\u8005\u4e4b\u95f4\u5b58\u5728\u516b\u500d\u5dee\u8ddd\u3002\n- glm-5.2 \u6ce2\u52a8\u6027\u6700\u9ad8\uff08\u53d8\u5f02\u7cfb\u6570 69.8%\uff09\uff0c\u5728 14.83 \u81f3 195.72 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\uff0c\u5176\u4e2d\u5305\u62ec 06:00 \u65f6\u4f4e\u81f3 14.83 \u8bcd\u5143/\u79d2\uff1bdeepseek-v4.1-flash \u4e5f\u4ece 06:40 \u7684 224.58 \u8bcd\u5143/\u79d2\u964d\u81f3 06:20 \u7684 35.29 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u6062\u590d\u3002\n- \u4e0d\u5b58\u5728\u7f3a\u5931\u6570\u636e\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62a5\u544a 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\u300196 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5e76\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u5b9e\u73b0 100.0% \u7684\u8986\u76d6\u7387\u3002","total_duration_ns":4185864512,"translated_at":"2026-10-07T07:03:08.534556+00:00","translation_eval_count":202,"translation_model":"glm-5.3","translation_prompt_eval_count":343,"translation_total_duration_ns":2301540577,"translation_wall_seconds":2.446,"valid_point_count":96,"wall_seconds":4.545},{"coverage_pct":100.0,"eval_count":311,"generated_at":"2026-10-07T06:03:07.525599+00:00","generated_label":"Oct 07, 2026 \u00b7 06:03 UTC","id":1190,"model":"glm-5.3","period_end":"2026-10-07T06:03:01+00:00","period_label":"Oct 07, 2026 \u00b7 02:03 UTC to Oct 07, 2026 \u00b7 06:03 UTC","period_start":"2026-10-07T02:03:01+00:00","prompt_eval_count":2250,"summary":"- deepseek-v4.1-flash is the strongest model at 171.8 token/s average throughput (124.2\u2013234.58 token/s range), while nemotron-3-ultra is the weakest at 21.53 token/s average, peaking at only 40.16 token/s.\n- glm-5.2 shows the most operationally significant volatility, with a 71.1% coefficient of variation and swings between 195.72 token/s at 04:00 and 14.83 token/s at 06:00, a -43.8% trend; glm-5.3 also oscillates between 67.57 and 182.91 token/s.\n- No missing-data limitation exists: all eight models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 171.8 token/s average throughput (124.2\u2013234.58 token/s range), while nemotron-3-ultra is the weakest at 21.53 token/s average, peaking at only 40.16 token/s.\n- glm-5.2 shows the most operationally significant volatility, with a 71.1% coefficient of variation and swings between 195.72 token/s at 04:00 and 14.83 token/s at 06:00, a -43.8% trend; glm-5.3 also oscillates between 67.57 and 182.91 token/s.\n- No missing-data limitation exists: all eight models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 171.8 token/s average throughput (124.2\u2013234.58 token/s range), while nemotron-3-ultra is the weakest at 21.53 token/s average, peaking at only 40.16 token/s.","glm-5.2 shows the most operationally significant volatility, with a 71.1% coefficient of variation and swings between 195.72 token/s at 04:00 and 14.83 token/s at 06:00, a -43.8% trend; glm-5.3 also oscillates between 67.57 and 182.91 token/s.","No missing-data limitation exists: all eight models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 171.8 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4\u4e3a 124.2\u2013234.58 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u662f\u6700\u5f31\u7684\u6a21\u578b\uff0c\u5e73\u5747\u4e3a 21.53 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u4ec5\u4e3a 40.16 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u5bf9\u8fd0\u8425\u5f71\u54cd\u6700\u5927\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 71.1%\uff0c\u5728 04:00 \u7684 195.72 \u8bcd\u5143/\u79d2\u4e0e 06:00 \u7684 14.83 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6446\u52a8\uff0c\u8d8b\u52bf\u4e3a -43.8%\uff1bglm-5.3 \u4e5f\u5728 67.57 \u4e0e 182.91 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u9707\u8361\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62e5\u6709 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u5171\u6709 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":3069001020,"translated_at":"2026-10-07T06:03:07.525599+00:00","translation_eval_count":207,"translation_model":"glm-5.3","translation_prompt_eval_count":355,"translation_total_duration_ns":2229575748,"translation_wall_seconds":2.372,"valid_point_count":96,"wall_seconds":3.239},{"coverage_pct":100.0,"eval_count":531,"generated_at":"2026-10-07T05:03:06.025022+00:00","generated_label":"Oct 07, 2026 \u00b7 05:03 UTC","id":1189,"model":"glm-5.3","period_end":"2026-10-07T05:03:01+00:00","period_label":"Oct 07, 2026 \u00b7 01:03 UTC to Oct 07, 2026 \u00b7 05:03 UTC","period_start":"2026-10-07T01:03:01+00:00","prompt_eval_count":2251,"summary":"- deepseek-v4.1-flash is the strongest model at 157.58 token/s average throughput, peaking at 234.58 token/s; nemotron-3-ultra is the weakest at 21.26 token/s average, never exceeding 37.98 token/s.\n- glm-5.2 shows the most volatility, with a 71.6% coefficient of variation and swings from 18.11 to 195.72 token/s, including a 195.72 token/s spike at 04:00 followed by a drop to 18.11 token/s at 04:20; deepseek-v4.1-flash also trended up 66.9% over the window.\n- No missing-data limitation exists: all eight models report 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 157.58 token/s average throughput, peaking at 234.58 token/s; nemotron-3-ultra is the weakest at 21.26 token/s average, never exceeding 37.98 token/s.\n- glm-5.2 shows the most volatility, with a 71.6% coefficient of variation and swings from 18.11 to 195.72 token/s, including a 195.72 token/s spike at 04:00 followed by a drop to 18.11 token/s at 04:20; deepseek-v4.1-flash also trended up 66.9% over the window.\n- No missing-data limitation exists: all eight models report 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 157.58 token/s average throughput, peaking at 234.58 token/s; nemotron-3-ultra is the weakest at 21.26 token/s average, never exceeding 37.98 token/s.","glm-5.2 shows the most volatility, with a 71.6% coefficient of variation and swings from 18.11 to 195.72 token/s, including a 195.72 token/s spike at 04:00 followed by a drop to 18.11 token/s at 04:20; deepseek-v4.1-flash also trended up 66.9% over the window.","No missing-data limitation exists: all eight models report 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 157.58 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 234.58 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 21.26 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 37.98 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u6ce2\u52a8\u6027\u6700\u5927\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 71.6%\uff0c\u6ce2\u52a8\u8303\u56f4\u4ece 18.11 \u5230 195.72 \u8bcd\u5143/\u79d2\uff0c\u5305\u62ec\u5728 04:00 \u51fa\u73b0 195.72 \u8bcd\u5143/\u79d2\u7684\u5cf0\u503c\uff0c\u968f\u540e\u5728 04:20 \u9aa4\u964d\u81f3 18.11 \u8bcd\u5143/\u79d2\uff1bdeepseek-v4.1-flash \u5728\u8be5\u65f6\u95f4\u7a97\u53e3\u5185\u4e5f\u5448 66.9% \u7684\u4e0a\u5347\u8d8b\u52bf\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62a5\u544a 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0cvalid_point_count \u4e3a 96\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":3155828023,"translated_at":"2026-10-07T05:03:06.025022+00:00","translation_eval_count":220,"translation_model":"glm-5.3","translation_prompt_eval_count":360,"translation_total_duration_ns":1036806875,"translation_wall_seconds":1.18,"valid_point_count":96,"wall_seconds":3.339},{"coverage_pct":100.0,"eval_count":477,"generated_at":"2026-10-07T04:03:06.069286+00:00","generated_label":"Oct 07, 2026 \u00b7 04:03 UTC","id":1188,"model":"glm-5.3","period_end":"2026-10-07T04:03:01+00:00","period_label":"Oct 07, 2026 \u00b7 00:03 UTC to Oct 07, 2026 \u00b7 04:03 UTC","period_start":"2026-10-07T00:03:01+00:00","prompt_eval_count":2248,"summary":"- deepseek-v4.1-flash is the strongest model by average throughput at 161.69 token/s (p95 228.39 token/s), while nemotron-3-ultra is the weakest at 24.56 token/s average, peaking at only 52.47 token/s.\n- The most operationally significant movement is gemma4:31b's steady decline from 159.5 token/s at 00:20 to 42.72 token/s at 04:00, a 35% downward trend, with its final three samples at 73.48, 103.13, 56.64, and 42.72 token/s. glm-5.2 also shows extreme volatility, ranging 13.44 to 195.72 token/s with a 70.5% coefficient of variation.\n- No missing-data limitation exists: all eight models have 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model by average throughput at 161.69 token/s (p95 228.39 token/s), while nemotron-3-ultra is the weakest at 24.56 token/s average, peaking at only 52.47 token/s.\n- The most operationally significant movement is gemma4:31b's steady decline from 159.5 token/s at 00:20 to 42.72 token/s at 04:00, a 35% downward trend, with its final three samples at 73.48, 103.13, 56.64, and 42.72 token/s. glm-5.2 also shows extreme volatility, ranging 13.44 to 195.72 token/s with a 70.5% coefficient of variation.\n- No missing-data limitation exists: all eight models have 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model by average throughput at 161.69 token/s (p95 228.39 token/s), while nemotron-3-ultra is the weakest at 24.56 token/s average, peaking at only 52.47 token/s.","The most operationally significant movement is gemma4:31b's steady decline from 159.5 token/s at 00:20 to 42.72 token/s at 04:00, a 35% downward trend, with its final three samples at 73.48, 103.13, 56.64, and 42.72 token/s. glm-5.2 also shows extreme volatility, ranging 13.44 to 195.72 token/s with a 70.5% coefficient of variation.","No missing-data limitation exists: all eight models have 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u8fbe 161.69 \u8bcd\u5143/\u79d2\uff08p95 \u4e3a 228.39 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4ec5 24.56 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u4e5f\u4ec5\u4e3a 52.47 \u8bcd\u5143/\u79d2\u3002\n- \u8fd0\u7ef4\u5c42\u9762\u6700\u663e\u8457\u7684\u53d8\u5316\u662f gemma4:31b \u4ece 00:20 \u7684 159.5 \u8bcd\u5143/\u79d2\u6301\u7eed\u4e0b\u964d\u81f3 04:00 \u7684 42.72 \u8bcd\u5143/\u79d2\uff0c\u5448 35% \u7684\u4e0b\u884c\u8d8b\u52bf\uff0c\u5176\u6700\u540e\u4e09\u4e2a\u91c7\u6837\u70b9\u5206\u522b\u4e3a 73.48\u3001103.13\u300156.64 \u548c 42.72 \u8bcd\u5143/\u79d2\u3002glm-5.2 \u540c\u6837\u8868\u73b0\u51fa\u6781\u5927\u7684\u6ce2\u52a8\u6027\uff0c\u8303\u56f4\u5728 13.44 \u81f3 195.72 \u8bcd\u5143/\u79d2\u4e4b\u95f4\uff0c\u53d8\u5f02\u7cfb\u6570\u8fbe 70.5%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u5747\u83b7\u5f97\u4e86 12 \u4e2a\u9884\u671f\u91c7\u6837\u4e2d\u7684 12 \u4e2a\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":2199895586,"translated_at":"2026-10-07T04:03:06.069286+00:00","translation_eval_count":343,"translation_model":"glm-5.3","translation_prompt_eval_count":384,"translation_total_duration_ns":1597958958,"translation_wall_seconds":1.746,"valid_point_count":96,"wall_seconds":2.368},{"coverage_pct":100.0,"eval_count":572,"generated_at":"2026-10-07T02:03:05.430352+00:00","generated_label":"Oct 07, 2026 \u00b7 02:03 UTC","id":1187,"model":"glm-5.3","period_end":"2026-10-07T02:03:01+00:00","period_label":"Oct 06, 2026 \u00b7 22:03 UTC to Oct 07, 2026 \u00b7 02:03 UTC","period_start":"2026-10-06T22:03:01+00:00","prompt_eval_count":2246,"summary":"- Strongest average throughput is deepseek-v4.1-flash at 163.39 token/s (p95 226.74 token/s); weakest is nemotron-3-ultra at 26.61 token/s (p95 47.12 token/s), roughly six times slower on average.\n- Most operationally significant volatility: deepseek-v4.1-flash fell from 219.23 token/s at 01:00 to 40.8 token/s at 01:40 before recovering to 172.35 token/s; glm-5.3 also swung between 182.85 and 60.98 token/s (cv 41.3%).\n- No missing data: all eight models report 12 of 12 samples, 96 valid points, 100.0% coverage; the limitation is only four hours of observations, so brief dips may not reflect steady-state throughput.","summary_en":"- Strongest average throughput is deepseek-v4.1-flash at 163.39 token/s (p95 226.74 token/s); weakest is nemotron-3-ultra at 26.61 token/s (p95 47.12 token/s), roughly six times slower on average.\n- Most operationally significant volatility: deepseek-v4.1-flash fell from 219.23 token/s at 01:00 to 40.8 token/s at 01:40 before recovering to 172.35 token/s; glm-5.3 also swung between 182.85 and 60.98 token/s (cv 41.3%).\n- No missing data: all eight models report 12 of 12 samples, 96 valid points, 100.0% coverage; the limitation is only four hours of observations, so brief dips may not reflect steady-state throughput.","summary_items":["Strongest average throughput is deepseek-v4.1-flash at 163.39 token/s (p95 226.74 token/s); weakest is nemotron-3-ultra at 26.61 token/s (p95 47.12 token/s), roughly six times slower on average.","Most operationally significant volatility: deepseek-v4.1-flash fell from 219.23 token/s at 01:00 to 40.8 token/s at 01:40 before recovering to 172.35 token/s; glm-5.3 also swung between 182.85 and 60.98 token/s (cv 41.3%).","No missing data: all eight models report 12 of 12 samples, 96 valid points, 100.0% coverage; the limitation is only four hours of observations, so brief dips may not reflect steady-state throughput."],"summary_zh":"- \u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u662f deepseek-v4.1-flash\uff0c\u8fbe 163.39 \u8bcd\u5143/\u79d2\uff08p95 \u4e3a 226.74 \u8bcd\u5143/\u79d2\uff09\uff1b\u6700\u5f31\u7684\u662f nemotron-3-ultra\uff0c\u4e3a 26.61 \u8bcd\u5143/\u79d2\uff08p95 \u4e3a 47.12 \u8bcd\u5143/\u79d2\uff09\uff0c\u5e73\u5747\u901f\u5ea6\u5927\u7ea6\u6162\u516d\u500d\u3002\n- \u8fd0\u8425\u5c42\u9762\u6700\u663e\u8457\u7684\u6ce2\u52a8\uff1adeepseek-v4.1-flash \u4ece 01:00 \u7684 219.23 \u8bcd\u5143/\u79d2\u964d\u81f3 01:40 \u7684 40.8 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u56de\u5347\u81f3 172.35 \u8bcd\u5143/\u79d2\uff1bglm-5.3 \u4e5f\u5728 182.85 \u548c 60.98 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u5927\u5e45\u6ce2\u52a8\uff08cv 41.3%\uff09\u3002\n- \u65e0\u7f3a\u5931\u6570\u636e\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62a5\u544a 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c100.0% \u7684\u8986\u76d6\u7387\uff1b\u5c40\u9650\u5728\u4e8e\u89c2\u6d4b\u65f6\u95f4\u4ec5\u6709\u56db\u5c0f\u65f6\uff0c\u56e0\u6b64\u77ed\u6682\u7684\u4e0b\u964d\u53ef\u80fd\u65e0\u6cd5\u53cd\u6620\u7a33\u6001\u541e\u5410\u91cf\u3002","total_duration_ns":2391005212,"translated_at":"2026-10-07T02:03:05.430352+00:00","translation_eval_count":233,"translation_model":"glm-5.3","translation_prompt_eval_count":361,"translation_total_duration_ns":938048626,"translation_wall_seconds":1.082,"valid_point_count":96,"wall_seconds":2.553},{"coverage_pct":100.0,"eval_count":736,"generated_at":"2026-10-07T01:03:10.223874+00:00","generated_label":"Oct 07, 2026 \u00b7 01:03 UTC","id":1186,"model":"glm-5.3","period_end":"2026-10-07T01:03:01+00:00","period_label":"Oct 06, 2026 \u00b7 21:03 UTC to Oct 07, 2026 \u00b7 01:03 UTC","period_start":"2026-10-06T21:03:01+00:00","prompt_eval_count":2249,"summary":"- deepseek-v4.1-flash is the strongest model at an average 184.46 token/s (peaking at 235.93 token/s), while nemotron-3-ultra is the weakest at an average 25.54 token/s, never exceeding 52.47 token/s.\n- glm-5.2 shows the most operationally significant volatility, with a 64.2% coefficient of variation and swings from 10.78 to 96.62 token/s within the four hours; nemotron-3-ultra is similarly erratic at 57.9%.\n- No missing-data limitation exists: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the period.","summary_en":"- deepseek-v4.1-flash is the strongest model at an average 184.46 token/s (peaking at 235.93 token/s), while nemotron-3-ultra is the weakest at an average 25.54 token/s, never exceeding 52.47 token/s.\n- glm-5.2 shows the most operationally significant volatility, with a 64.2% coefficient of variation and swings from 10.78 to 96.62 token/s within the four hours; nemotron-3-ultra is similarly erratic at 57.9%.\n- No missing-data limitation exists: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the period.","summary_items":["deepseek-v4.1-flash is the strongest model at an average 184.46 token/s (peaking at 235.93 token/s), while nemotron-3-ultra is the weakest at an average 25.54 token/s, never exceeding 52.47 token/s.","glm-5.2 shows the most operationally significant volatility, with a 64.2% coefficient of variation and swings from 10.78 to 96.62 token/s within the four hours; nemotron-3-ultra is similarly erratic at 57.9%.","No missing-data limitation exists: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the period."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u901f\u5ea6\u4e3a 184.46 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c\u8fbe 235.93 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u901f\u5ea6\u4ec5\u4e3a 25.54 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 52.47 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u5bf9\u8fd0\u884c\u5f71\u54cd\u6700\u5927\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 64.2%\uff0c\u5728\u56db\u5c0f\u65f6\u5185\u901f\u5ea6\u5728 10.78 \u81f3 96.62 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6446\u52a8\uff1bnemotron-3-ultra \u540c\u6837\u4e0d\u7a33\u5b9a\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 57.9%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62a5\u544a\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u6574\u4e2a\u671f\u95f4\u5171\u6709 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":6111008400,"translated_at":"2026-10-07T01:03:10.223874+00:00","translation_eval_count":186,"translation_model":"glm-5.3","translation_prompt_eval_count":329,"translation_total_duration_ns":2190872924,"translation_wall_seconds":2.338,"valid_point_count":96,"wall_seconds":6.273},{"coverage_pct":100.0,"eval_count":801,"generated_at":"2026-10-07T00:03:08.120779+00:00","generated_label":"Oct 07, 2026 \u00b7 00:03 UTC","id":1185,"model":"glm-5.3","period_end":"2026-10-07T00:03:02+00:00","period_label":"Oct 06, 2026 \u00b7 20:03 UTC to Oct 07, 2026 \u00b7 00:03 UTC","period_start":"2026-10-06T20:03:02+00:00","prompt_eval_count":2246,"summary":"- deepseek-v4.1-flash is the strongest model at 186.75 token/s average throughput (peak 235.93 token/s), while nemotron-3-ultra is the weakest at 22.48 token/s average (minimum 3.15 token/s).\n- glm-5.2 shows the most extreme volatility, ranging from 10.78 to 216.7 token/s with a 94.6% coefficient of variation; gemma4:31b also swung from 146.13 down to 43.95 token/s late in the window.\n- No missing-data limitation exists: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 186.75 token/s average throughput (peak 235.93 token/s), while nemotron-3-ultra is the weakest at 22.48 token/s average (minimum 3.15 token/s).\n- glm-5.2 shows the most extreme volatility, ranging from 10.78 to 216.7 token/s with a 94.6% coefficient of variation; gemma4:31b also swung from 146.13 down to 43.95 token/s late in the window.\n- No missing-data limitation exists: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 186.75 token/s average throughput (peak 235.93 token/s), while nemotron-3-ultra is the weakest at 22.48 token/s average (minimum 3.15 token/s).","glm-5.2 shows the most extreme volatility, ranging from 10.78 to 216.7 token/s with a 94.6% coefficient of variation; gemma4:31b also swung from 146.13 down to 43.95 token/s late in the window.","No missing-data limitation exists: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 186.75 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 235.93 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 22.48 \u8bcd\u5143/\u79d2\uff08\u6700\u4f4e 3.15 \u8bcd\u5143/\u79d2\uff09\u3002\n- glm-5.2 \u6ce2\u52a8\u6700\u4e3a\u5267\u70c8\uff0c\u8303\u56f4\u4ece 10.78 \u5230 216.7 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u8fbe 94.6%\uff1bgemma4:31b \u5728\u65f6\u95f4\u7a97\u53e3\u540e\u671f\u4e5f\u4ece 146.13 \u9aa4\u964d\u81f3 43.95 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62a5\u544a\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u8986\u76d6\u7387\u8fbe\u5230 100.0%\u3002","total_duration_ns":3414989203,"translated_at":"2026-10-07T00:03:08.120779+00:00","translation_eval_count":178,"translation_model":"glm-5.3","translation_prompt_eval_count":330,"translation_total_duration_ns":2301901109,"translation_wall_seconds":2.448,"valid_point_count":96,"wall_seconds":3.581},{"coverage_pct":100.0,"eval_count":321,"generated_at":"2026-10-06T23:03:05.067977+00:00","generated_label":"Oct 06, 2026 \u00b7 23:03 UTC","id":1184,"model":"glm-5.3","period_end":"2026-10-06T23:03:02+00:00","period_label":"Oct 06, 2026 \u00b7 19:03 UTC to Oct 06, 2026 \u00b7 23:03 UTC","period_start":"2026-10-06T19:03:02+00:00","prompt_eval_count":2248,"summary":"- Strongest average throughput was deepseek-v4.1-flash at 193.04 token/s (peaking at 235.93 token/s), while nemotron-3-ultra was weakest at 21.48 token/s, never exceeding 42.75 token/s.\n- glm-5.2 showed the most extreme volatility, with a coefficient of variation of 96.7%: it spiked to 216.7 token/s at 20:40 then fell to 10.78 token/s at 21:20, and all models trended downward except deepseek-v4-pro (+3.9%) and nemotron-3-ultra (+27.3%).\n- No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, and the dataset achieved 100.0% coverage with 96 of 96 valid points.","summary_en":"- Strongest average throughput was deepseek-v4.1-flash at 193.04 token/s (peaking at 235.93 token/s), while nemotron-3-ultra was weakest at 21.48 token/s, never exceeding 42.75 token/s.\n- glm-5.2 showed the most extreme volatility, with a coefficient of variation of 96.7%: it spiked to 216.7 token/s at 20:40 then fell to 10.78 token/s at 21:20, and all models trended downward except deepseek-v4-pro (+3.9%) and nemotron-3-ultra (+27.3%).\n- No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, and the dataset achieved 100.0% coverage with 96 of 96 valid points.","summary_items":["Strongest average throughput was deepseek-v4.1-flash at 193.04 token/s (peaking at 235.93 token/s), while nemotron-3-ultra was weakest at 21.48 token/s, never exceeding 42.75 token/s.","glm-5.2 showed the most extreme volatility, with a coefficient of variation of 96.7%: it spiked to 216.7 token/s at 20:40 then fell to 10.78 token/s at 21:20, and all models trended downward except deepseek-v4-pro (+3.9%) and nemotron-3-ultra (+27.3%).","No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, and the dataset achieved 100.0% coverage with 96 of 96 valid points."],"summary_zh":"- \u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u662f deepseek-v4.1-flash\uff0c\u8fbe 193.04 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 235.93 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4ec5 21.48 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 42.75 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u7684\u6ce2\u52a8\u6027\u6700\u4e3a\u6781\u7aef\uff0c\u53d8\u5f02\u7cfb\u6570\u8fbe 96.7%\uff1a\u5b83\u5728 20:40 \u98d9\u5347\u81f3 216.7 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u5728 21:20 \u8dcc\u81f3 10.78 \u8bcd\u5143/\u79d2\uff1b\u9664 deepseek-v4-pro\uff08+3.9%\uff09\u548c nemotron-3-ultra\uff08+27.3%\uff09\u5916\uff0c\u6240\u6709\u6a21\u578b\u5747\u5448\u4e0b\u964d\u8d8b\u52bf\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8 8 \u4e2a\u6a21\u578b\u5747\u4ea4\u4ed8\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u6570\u636e\u96c6\u5b9e\u73b0\u4e86 100.0% \u7684\u8986\u76d6\u7387\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\u4e2d\u7684 96 \u4e2a\u3002","total_duration_ns":1521132394,"translated_at":"2026-10-06T23:03:05.067977+00:00","translation_eval_count":220,"translation_model":"glm-5.3","translation_prompt_eval_count":354,"translation_total_duration_ns":989604429,"translation_wall_seconds":1.136,"valid_point_count":96,"wall_seconds":1.911},{"coverage_pct":100.0,"eval_count":538,"generated_at":"2026-10-06T22:03:06.297758+00:00","generated_label":"Oct 06, 2026 \u00b7 22:03 UTC","id":1183,"model":"glm-5.3","period_end":"2026-10-06T22:03:01+00:00","period_label":"Oct 06, 2026 \u00b7 18:03 UTC to Oct 06, 2026 \u00b7 22:03 UTC","period_start":"2026-10-06T18:03:01+00:00","prompt_eval_count":2249,"summary":"- deepseek-v4.1-flash is the strongest model at 186.07 token/s average throughput, peaking at 230.07 token/s, while nemotron-3-ultra is the weakest at 20.1 token/s average, with a maximum of only 40.72 token/s.\n- glm-5.2 shows the most operationally significant volatility, ranging from 10.78 to 216.7 token/s with a 105.7% coefficient of variation, including a spike to 216.7 token/s at 20:40 UTC followed by a drop to 10.78 token/s at 21:20 UTC; deepseek-v4.1-flash also trended upward 9.3%.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 186.07 token/s average throughput, peaking at 230.07 token/s, while nemotron-3-ultra is the weakest at 20.1 token/s average, with a maximum of only 40.72 token/s.\n- glm-5.2 shows the most operationally significant volatility, ranging from 10.78 to 216.7 token/s with a 105.7% coefficient of variation, including a spike to 216.7 token/s at 20:40 UTC followed by a drop to 10.78 token/s at 21:20 UTC; deepseek-v4.1-flash also trended upward 9.3%.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 186.07 token/s average throughput, peaking at 230.07 token/s, while nemotron-3-ultra is the weakest at 20.1 token/s average, with a maximum of only 40.72 token/s.","glm-5.2 shows the most operationally significant volatility, ranging from 10.78 to 216.7 token/s with a 105.7% coefficient of variation, including a spike to 216.7 token/s at 20:40 UTC followed by a drop to 10.78 token/s at 21:20 UTC; deepseek-v4.1-flash also trended upward 9.3%.","No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 186.07 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 230.07 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 20.1 \u8bcd\u5143/\u79d2\uff0c\u6700\u5927\u503c\u4ec5\u4e3a 40.72 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u5bf9\u8fd0\u8425\u5f71\u54cd\u6700\u4e3a\u663e\u8457\u7684\u6ce2\u52a8\u6027\uff0c\u8303\u56f4\u4ece 10.78 \u5230 216.7 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 105.7%\uff0c\u5176\u4e2d\u5305\u62ec\u5728 UTC 20:40 \u98d9\u5347\u81f3 216.7 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u5728 UTC 21:20 \u9aa4\u964d\u81f3 10.78 \u8bcd\u5143/\u79d2\uff1bdeepseek-v4.1-flash \u4e5f\u5448 9.3% \u7684\u4e0a\u5347\u8d8b\u52bf\u3002\n- \u4e0d\u5b58\u5728\u7f3a\u5931\u6570\u636e\u7684\u9650\u5236\uff1a\u6240\u6709\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0cvalid_point_count \u4e3a 96\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":2703646988,"translated_at":"2026-10-06T22:03:06.297758+00:00","translation_eval_count":453,"translation_model":"glm-5.3","translation_prompt_eval_count":365,"translation_total_duration_ns":1908723482,"translation_wall_seconds":2.051,"valid_point_count":96,"wall_seconds":2.862},{"coverage_pct":100.0,"eval_count":274,"generated_at":"2026-10-06T21:03:07.272771+00:00","generated_label":"Oct 06, 2026 \u00b7 21:03 UTC","id":1182,"model":"glm-5.3","period_end":"2026-10-06T21:03:01+00:00","period_label":"Oct 06, 2026 \u00b7 17:03 UTC to Oct 06, 2026 \u00b7 21:03 UTC","period_start":"2026-10-06T17:03:01+00:00","prompt_eval_count":2246,"summary":"- deepseek-v4.1-flash is the strongest model at 184.08 token/s average throughput, peaking at 230.07 token/s, while nemotron-3-ultra is the weakest at 21.36 token/s average, dipping to 4.27 token/s.\n- glm-5.2 shows the most extreme volatility, ranging from 11.04 to 216.7 token/s with a 96.0% coefficient of variation; glm-5.3 posted the steepest climb, up 52.0% to a 175.32 token/s peak at 20:20.\n- No missing-data limitation exists in this window: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage.","summary_en":"- deepseek-v4.1-flash is the strongest model at 184.08 token/s average throughput, peaking at 230.07 token/s, while nemotron-3-ultra is the weakest at 21.36 token/s average, dipping to 4.27 token/s.\n- glm-5.2 shows the most extreme volatility, ranging from 11.04 to 216.7 token/s with a 96.0% coefficient of variation; glm-5.3 posted the steepest climb, up 52.0% to a 175.32 token/s peak at 20:20.\n- No missing-data limitation exists in this window: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage.","summary_items":["deepseek-v4.1-flash is the strongest model at 184.08 token/s average throughput, peaking at 230.07 token/s, while nemotron-3-ultra is the weakest at 21.36 token/s average, dipping to 4.27 token/s.","glm-5.2 shows the most extreme volatility, ranging from 11.04 to 216.7 token/s with a 96.0% coefficient of variation; glm-5.3 posted the steepest climb, up 52.0% to a 175.32 token/s peak at 20:20.","No missing-data limitation exists in this window: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 184.08 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 230.07 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 21.36 \u8bcd\u5143/\u79d2\uff0c\u6700\u4f4e\u964d\u81f3 4.27 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u6ce2\u52a8\u6700\u4e3a\u5267\u70c8\uff0c\u8303\u56f4\u4ece 11.04 \u5230 216.7 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 96.0%\uff1bglm-5.3 \u7684\u6500\u5347\u6700\u4e3a\u9661\u5ced\uff0c\u4e0a\u5347 52.0%\uff0c\u5728 20:20 \u8fbe\u5230 175.32 \u8bcd\u5143/\u79d2\u7684\u5cf0\u503c\u3002\n- \u672c\u65f6\u95f4\u7a97\u53e3\u5185\u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u4ea4\u4ed8\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5171\u6709 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":2985790951,"translated_at":"2026-10-06T21:03:07.272771+00:00","translation_eval_count":184,"translation_model":"glm-5.3","translation_prompt_eval_count":338,"translation_total_duration_ns":2462529294,"translation_wall_seconds":2.611,"valid_point_count":96,"wall_seconds":3.155},{"coverage_pct":100.0,"eval_count":513,"generated_at":"2026-10-06T20:03:06.427742+00:00","generated_label":"Oct 06, 2026 \u00b7 20:03 UTC","id":1181,"model":"glm-5.3","period_end":"2026-10-06T20:03:02+00:00","period_label":"Oct 06, 2026 \u00b7 16:03 UTC to Oct 06, 2026 \u00b7 20:03 UTC","period_start":"2026-10-06T16:03:02+00:00","prompt_eval_count":2246,"summary":"- deepseek-v4.1-flash is the strongest model at 181.27 token/s average throughput (peak 228.81 token/s), while nemotron-3-ultra is the weakest at 20.47 token/s average, never exceeding 40.72 token/s.\n- glm-5.2 shows the most operationally significant volatility, with a 70.2% coefficient of variation and swings from 13.18 to 134.30 token/s; glm-5.3 similarly ranged from 16.72 to 163.17 token/s, indicating unstable throughput for both.\n- No missing data: all eight models have 12 of 12 samples with 100.0% coverage and 96 valid points, so the only limitation is the short four-hour window, which may not represent longer-term behavior.","summary_en":"- deepseek-v4.1-flash is the strongest model at 181.27 token/s average throughput (peak 228.81 token/s), while nemotron-3-ultra is the weakest at 20.47 token/s average, never exceeding 40.72 token/s.\n- glm-5.2 shows the most operationally significant volatility, with a 70.2% coefficient of variation and swings from 13.18 to 134.30 token/s; glm-5.3 similarly ranged from 16.72 to 163.17 token/s, indicating unstable throughput for both.\n- No missing data: all eight models have 12 of 12 samples with 100.0% coverage and 96 valid points, so the only limitation is the short four-hour window, which may not represent longer-term behavior.","summary_items":["deepseek-v4.1-flash is the strongest model at 181.27 token/s average throughput (peak 228.81 token/s), while nemotron-3-ultra is the weakest at 20.47 token/s average, never exceeding 40.72 token/s.","glm-5.2 shows the most operationally significant volatility, with a 70.2% coefficient of variation and swings from 13.18 to 134.30 token/s; glm-5.3 similarly ranged from 16.72 to 163.17 token/s, indicating unstable throughput for both.","No missing data: all eight models have 12 of 12 samples with 100.0% coverage and 96 valid points, so the only limitation is the short four-hour window, which may not represent longer-term behavior."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 181.27 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 228.81 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4ec5\u4e3a 20.47 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 40.72 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u5bf9\u8fd0\u884c\u5f71\u54cd\u6700\u663e\u8457\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 70.2%\uff0c\u6ce2\u52a8\u8303\u56f4\u4ece 13.18 \u5230 134.30 \u8bcd\u5143/\u79d2\uff1bglm-5.3 \u7684\u8303\u56f4\u540c\u6837\u4e3a 16.72 \u5230 163.17 \u8bcd\u5143/\u79d2\uff0c\u8868\u660e\u4e24\u8005\u7684\u541e\u5410\u91cf\u5747\u4e0d\u7a33\u5b9a\u3002\n- \u65e0\u7f3a\u5931\u6570\u636e\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u56e0\u6b64\u552f\u4e00\u7684\u5c40\u9650\u662f\u4ec5\u56db\u5c0f\u65f6\u7684\u77ed\u65f6\u95f4\u7a97\u53e3\uff0c\u53ef\u80fd\u65e0\u6cd5\u4ee3\u8868\u957f\u671f\u8868\u73b0\u3002","total_duration_ns":2612858808,"translated_at":"2026-10-06T20:03:06.427742+00:00","translation_eval_count":213,"translation_model":"glm-5.3","translation_prompt_eval_count":344,"translation_total_duration_ns":1095292527,"translation_wall_seconds":1.415,"valid_point_count":96,"wall_seconds":2.953},{"coverage_pct":100.0,"eval_count":483,"generated_at":"2026-10-06T19:03:09.630613+00:00","generated_label":"Oct 06, 2026 \u00b7 19:03 UTC","id":1180,"model":"glm-5.3","period_end":"2026-10-06T19:03:01+00:00","period_label":"Oct 06, 2026 \u00b7 15:03 UTC to Oct 06, 2026 \u00b7 19:03 UTC","period_start":"2026-10-06T15:03:01+00:00","prompt_eval_count":2245,"summary":"- deepseek-v4.1-flash is the strongest model at 164.93 token/s average throughput, peaking at 228.81 token/s; nemotron-3-ultra is the weakest at 20.23 token/s average, never exceeding 40.72 token/s.\n- glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 68.5% and swings between 13.18 and 134.3 token/s; glm-5.3 similarly ranged from 16.72 to 142.96 token/s, indicating unstable throughput.\n- No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 164.93 token/s average throughput, peaking at 228.81 token/s; nemotron-3-ultra is the weakest at 20.23 token/s average, never exceeding 40.72 token/s.\n- glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 68.5% and swings between 13.18 and 134.3 token/s; glm-5.3 similarly ranged from 16.72 to 142.96 token/s, indicating unstable throughput.\n- No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 164.93 token/s average throughput, peaking at 228.81 token/s; nemotron-3-ultra is the weakest at 20.23 token/s average, never exceeding 40.72 token/s.","glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 68.5% and swings between 13.18 and 134.3 token/s; glm-5.3 similarly ranged from 16.72 to 142.96 token/s, indicating unstable throughput.","No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 164.93 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 228.81 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 20.23 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 40.72 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u5bf9\u8fd0\u884c\u5f71\u54cd\u6700\u5927\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 68.5%\uff0c\u5728 13.18 \u81f3 134.3 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\uff1bglm-5.3 \u540c\u6837\u5728 16.72 \u81f3 142.96 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\uff0c\u8868\u660e\u541e\u5410\u91cf\u4e0d\u7a33\u5b9a\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u8bb0\u5f55\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u5171\u8ba1 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":3824506953,"translated_at":"2026-10-06T19:03:09.630613+00:00","translation_eval_count":185,"translation_model":"glm-5.3","translation_prompt_eval_count":334,"translation_total_duration_ns":3609883888,"translation_wall_seconds":3.754,"valid_point_count":96,"wall_seconds":3.989},{"coverage_pct":100.0,"eval_count":486,"generated_at":"2026-10-06T18:03:05.035653+00:00","generated_label":"Oct 06, 2026 \u00b7 18:03 UTC","id":1179,"model":"glm-5.3","period_end":"2026-10-06T18:03:01+00:00","period_label":"Oct 06, 2026 \u00b7 14:03 UTC to Oct 06, 2026 \u00b7 18:03 UTC","period_start":"2026-10-06T14:03:01+00:00","prompt_eval_count":2245,"summary":"- deepseek-v4.1-flash is the strongest model at 162.24 token/s average throughput (peak 228.81 token/s), while nemotron-3-ultra is the weakest at 19.68 token/s average (minimum 7.74 token/s).\n- glm-5.2 shows the most operationally significant volatility, with an 80.5% coefficient of variation, swings from 16.12 to 195.22 token/s, and a -44.6% trend; deepseek-v4.1-flash trended up 32.2%, closing at 185.94 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 162.24 token/s average throughput (peak 228.81 token/s), while nemotron-3-ultra is the weakest at 19.68 token/s average (minimum 7.74 token/s).\n- glm-5.2 shows the most operationally significant volatility, with an 80.5% coefficient of variation, swings from 16.12 to 195.22 token/s, and a -44.6% trend; deepseek-v4.1-flash trended up 32.2%, closing at 185.94 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 162.24 token/s average throughput (peak 228.81 token/s), while nemotron-3-ultra is the weakest at 19.68 token/s average (minimum 7.74 token/s).","glm-5.2 shows the most operationally significant volatility, with an 80.5% coefficient of variation, swings from 16.12 to 195.22 token/s, and a -44.6% trend; deepseek-v4.1-flash trended up 32.2%, closing at 185.94 token/s.","No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6027\u80fd\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 162.24 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 228.81 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6027\u80fd\u6700\u5f31\uff0c\u5e73\u5747\u4e3a 19.68 \u8bcd\u5143/\u79d2\uff08\u6700\u4f4e 7.74 \u8bcd\u5143/\u79d2\uff09\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u5bf9\u8fd0\u8425\u5f71\u54cd\u6700\u5927\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 80.5%\uff0c\u6ce2\u52a8\u8303\u56f4\u4ece 16.12 \u5230 195.22 \u8bcd\u5143/\u79d2\uff0c\u8d8b\u52bf\u4e3a -44.6%\uff1bdeepseek-v4.1-flash \u5448\u4e0a\u5347\u8d8b\u52bf\uff0c\u6da8\u5e45 32.2%\uff0c\u6536\u4e8e 185.94 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\u300196 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":2320856693,"translated_at":"2026-10-06T18:03:05.035653+00:00","translation_eval_count":198,"translation_model":"glm-5.3","translation_prompt_eval_count":339,"translation_total_duration_ns":1047025081,"translation_wall_seconds":1.189,"valid_point_count":96,"wall_seconds":2.49},{"coverage_pct":100.0,"eval_count":866,"generated_at":"2026-10-06T17:03:06.770434+00:00","generated_label":"Oct 06, 2026 \u00b7 17:03 UTC","id":1178,"model":"glm-5.3","period_end":"2026-10-06T17:03:01+00:00","period_label":"Oct 06, 2026 \u00b7 13:03 UTC to Oct 06, 2026 \u00b7 17:03 UTC","period_start":"2026-10-06T13:03:01+00:00","prompt_eval_count":2248,"summary":"- deepseek-v4.1-flash is the strongest model at 166.42 token/s average throughput, peaking at 229.21 token/s; nemotron-3-ultra is weakest at 19.27 token/s average, never exceeding 33.31 token/s.\n- glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 80.0% and a swing from 16.12 to 195.22 token/s, plus a -42.2% trend; deepseek-v4.1-flash is steadier at 27.8% variation despite a 229.21-to-85.42 token/s range.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.","summary_en":"- deepseek-v4.1-flash is the strongest model at 166.42 token/s average throughput, peaking at 229.21 token/s; nemotron-3-ultra is weakest at 19.27 token/s average, never exceeding 33.31 token/s.\n- glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 80.0% and a swing from 16.12 to 195.22 token/s, plus a -42.2% trend; deepseek-v4.1-flash is steadier at 27.8% variation despite a 229.21-to-85.42 token/s range.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.","summary_items":["deepseek-v4.1-flash is the strongest model at 166.42 token/s average throughput, peaking at 229.21 token/s; nemotron-3-ultra is weakest at 19.27 token/s average, never exceeding 33.31 token/s.","glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 80.0% and a swing from 16.12 to 195.22 token/s, plus a -42.2% trend; deepseek-v4.1-flash is steadier at 27.8% variation despite a 229.21-to-85.42 token/s range.","No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 166.42 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 229.21 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 19.27 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 33.31 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u5bf9\u8fd0\u884c\u5f71\u54cd\u6700\u4e3a\u663e\u8457\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 80.0%\uff0c\u6ce2\u52a8\u8303\u56f4\u4e3a 16.12 \u81f3 195.22 \u8bcd\u5143/\u79d2\uff0c\u4e14\u8d8b\u52bf\u4e3a -42.2%\uff1bdeepseek-v4.1-flash \u5219\u66f4\u4e3a\u7a33\u5b9a\uff0c\u5c3d\u7ba1\u6ce2\u52a8\u8303\u56f4\u4e3a 229.21 \u81f3 85.42 \u8bcd\u5143/\u79d2\uff0c\u4f46\u53d8\u5f02\u4ec5\u4e3a 27.8%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0cvalid_point_count \u4e3a 96\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u56e0\u6b64\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5f97\u5230\u4e86\u5b8c\u6574\u5448\u73b0\u3002","total_duration_ns":4014518272,"translated_at":"2026-10-06T17:03:06.770434+00:00","translation_eval_count":211,"translation_model":"glm-5.3","translation_prompt_eval_count":358,"translation_total_duration_ns":1133964436,"translation_wall_seconds":1.288,"valid_point_count":96,"wall_seconds":4.19},{"coverage_pct":100.0,"eval_count":664,"generated_at":"2026-10-06T16:03:06.150380+00:00","generated_label":"Oct 06, 2026 \u00b7 16:03 UTC","id":1177,"model":"glm-5.3","period_end":"2026-10-06T16:03:01+00:00","period_label":"Oct 06, 2026 \u00b7 12:03 UTC to Oct 06, 2026 \u00b7 16:03 UTC","period_start":"2026-10-06T12:03:01+00:00","prompt_eval_count":2247,"summary":"- deepseek-v4.1-flash is the strongest model at 160.42 token/s average throughput (peak 229.21 token/s), while nemotron-3-ultra is the weakest at 18.01 token/s average (peak 33.31 token/s).\n- glm-5.2 shows the most volatility, with a coefficient of variation of 86.5% and swings from 7.84 to 195.22 token/s; a broad dip near 14:00 UTC hit nearly all models, including minimax-m3 at 4.59 token/s and gemma4:31b at 47.71 token/s.\n- No missing-data limitation exists: all eight models have 12 of 12 samples and 96 valid points at 100% coverage, though the four-hour window alone cannot confirm whether the 14:00 dip is recurring.","summary_en":"- deepseek-v4.1-flash is the strongest model at 160.42 token/s average throughput (peak 229.21 token/s), while nemotron-3-ultra is the weakest at 18.01 token/s average (peak 33.31 token/s).\n- glm-5.2 shows the most volatility, with a coefficient of variation of 86.5% and swings from 7.84 to 195.22 token/s; a broad dip near 14:00 UTC hit nearly all models, including minimax-m3 at 4.59 token/s and gemma4:31b at 47.71 token/s.\n- No missing-data limitation exists: all eight models have 12 of 12 samples and 96 valid points at 100% coverage, though the four-hour window alone cannot confirm whether the 14:00 dip is recurring.","summary_items":["deepseek-v4.1-flash is the strongest model at 160.42 token/s average throughput (peak 229.21 token/s), while nemotron-3-ultra is the weakest at 18.01 token/s average (peak 33.31 token/s).","glm-5.2 shows the most volatility, with a coefficient of variation of 86.5% and swings from 7.84 to 195.22 token/s; a broad dip near 14:00 UTC hit nearly all models, including minimax-m3 at 4.59 token/s and gemma4:31b at 47.71 token/s.","No missing-data limitation exists: all eight models have 12 of 12 samples and 96 valid points at 100% coverage, though the four-hour window alone cannot confirm whether the 14:00 dip is recurring."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 160.42 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 229.21 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 18.01 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 33.31 \u8bcd\u5143/\u79d2\uff09\u3002\n- glm-5.2 \u6ce2\u52a8\u6700\u5927\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 86.5%\uff0c\u5728 7.84 \u81f3 195.22 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\uff1bUTC \u65f6\u95f4 14:00 \u9644\u8fd1\u7684\u5927\u8303\u56f4\u4f4e\u8c37\u51e0\u4e4e\u5f71\u54cd\u4e86\u6240\u6709\u6a21\u578b\uff0c\u5305\u62ec minimax-m3 \u7684 4.59 \u8bcd\u5143/\u79d2\u548c gemma4:31b \u7684 47.71 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100%\uff0c\u4f46\u4ec5\u51ed\u8fd9\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u65e0\u6cd5\u786e\u8ba4 14:00 \u7684\u4f4e\u8c37\u662f\u5426\u53cd\u590d\u51fa\u73b0\u3002","total_duration_ns":3177846202,"translated_at":"2026-10-06T16:03:06.150380+00:00","translation_eval_count":218,"translation_model":"glm-5.3","translation_prompt_eval_count":356,"translation_total_duration_ns":1053745693,"translation_wall_seconds":1.2,"valid_point_count":96,"wall_seconds":3.514},{"coverage_pct":100.0,"eval_count":899,"generated_at":"2026-10-06T15:03:08.468306+00:00","generated_label":"Oct 06, 2026 \u00b7 15:03 UTC","id":1176,"model":"glm-5.3","period_end":"2026-10-06T15:03:02+00:00","period_label":"Oct 06, 2026 \u00b7 11:03 UTC to Oct 06, 2026 \u00b7 15:03 UTC","period_start":"2026-10-06T11:03:02+00:00","prompt_eval_count":2247,"summary":"- deepseek-v4.1-flash is the strongest model at 169.92 token/s average throughput (peak 229.21 token/s), while nemotron-3-ultra is the weakest at 19.01 token/s average (peak 33.31 token/s).\n- The 14:00 UTC observation shows a synchronized dip: gemma4:31b fell to 47.71 token/s, minimax-m3 to 4.59 token/s, and deepseek-v4-pro to 23.05 token/s; glm-5.2 is the most volatile model with a 79.8% coefficient of variation, swinging between 7.84 and 195.97 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage.","summary_en":"- deepseek-v4.1-flash is the strongest model at 169.92 token/s average throughput (peak 229.21 token/s), while nemotron-3-ultra is the weakest at 19.01 token/s average (peak 33.31 token/s).\n- The 14:00 UTC observation shows a synchronized dip: gemma4:31b fell to 47.71 token/s, minimax-m3 to 4.59 token/s, and deepseek-v4-pro to 23.05 token/s; glm-5.2 is the most volatile model with a 79.8% coefficient of variation, swinging between 7.84 and 195.97 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage.","summary_items":["deepseek-v4.1-flash is the strongest model at 169.92 token/s average throughput (peak 229.21 token/s), while nemotron-3-ultra is the weakest at 19.01 token/s average (peak 33.31 token/s).","The 14:00 UTC observation shows a synchronized dip: gemma4:31b fell to 47.71 token/s, minimax-m3 to 4.59 token/s, and deepseek-v4-pro to 23.05 token/s; glm-5.2 is the most volatile model with a 79.8% coefficient of variation, swinging between 7.84 and 195.97 token/s.","No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 169.92 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 229.21 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u662f\u6700\u5f31\u7684\u6a21\u578b\uff0c\u5e73\u5747\u4e3a 19.01 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 33.31 \u8bcd\u5143/\u79d2\uff09\u3002\n- 14:00 UTC \u7684\u89c2\u6d4b\u6570\u636e\u663e\u793a\u4e00\u6b21\u540c\u6b65\u4e0b\u8dcc\uff1agemma4:31b \u964d\u81f3 47.71 \u8bcd\u5143/\u79d2\uff0cminimax-m3 \u964d\u81f3 4.59 \u8bcd\u5143/\u79d2\uff0cdeepseek-v4-pro \u964d\u81f3 23.05 \u8bcd\u5143/\u79d2\uff1bglm-5.2 \u662f\u6ce2\u52a8\u6700\u5927\u7684\u6a21\u578b\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 79.8%\uff0c\u5728 7.84 \u81f3 195.97 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u6570\u636e\u96c6\u62a5\u544a\u4e86 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":4971084733,"translated_at":"2026-10-06T15:03:08.468306+00:00","translation_eval_count":209,"translation_model":"glm-5.3","translation_prompt_eval_count":355,"translation_total_duration_ns":1060236662,"translation_wall_seconds":1.206,"valid_point_count":96,"wall_seconds":5.144},{"coverage_pct":100.0,"eval_count":767,"generated_at":"2026-10-06T14:03:19.109019+00:00","generated_label":"Oct 06, 2026 \u00b7 14:03 UTC","id":1175,"model":"glm-5.3","period_end":"2026-10-06T14:03:01+00:00","period_label":"Oct 06, 2026 \u00b7 10:03 UTC to Oct 06, 2026 \u00b7 14:03 UTC","period_start":"2026-10-06T10:03:01+00:00","prompt_eval_count":2247,"summary":"- deepseek-v4.1-flash is the strongest model at 182.63 token/s average throughput (peak 230.12 token/s), while nemotron-3-ultra is the weakest at 18.48 token/s average (peak 33.31 token/s).\n- The most significant trend is minimax-m3's decline of 55.1% over the window, ending at 4.59 token/s; glm-5.2 shows the highest volatility, swinging between 4.76 and 195.97 token/s with an 85.0% coefficient of variation.\n- No missing-data limitation: all 8 models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage, so the four-hour window is complete.","summary_en":"- deepseek-v4.1-flash is the strongest model at 182.63 token/s average throughput (peak 230.12 token/s), while nemotron-3-ultra is the weakest at 18.48 token/s average (peak 33.31 token/s).\n- The most significant trend is minimax-m3's decline of 55.1% over the window, ending at 4.59 token/s; glm-5.2 shows the highest volatility, swinging between 4.76 and 195.97 token/s with an 85.0% coefficient of variation.\n- No missing-data limitation: all 8 models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage, so the four-hour window is complete.","summary_items":["deepseek-v4.1-flash is the strongest model at 182.63 token/s average throughput (peak 230.12 token/s), while nemotron-3-ultra is the weakest at 18.48 token/s average (peak 33.31 token/s).","The most significant trend is minimax-m3's decline of 55.1% over the window, ending at 4.59 token/s; glm-5.2 shows the highest volatility, swinging between 4.76 and 195.97 token/s with an 85.0% coefficient of variation.","No missing-data limitation: all 8 models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage, so the four-hour window is complete."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 182.63 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 230.12 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 18.48 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 33.31 \u8bcd\u5143/\u79d2\uff09\u3002\n- \u6700\u663e\u8457\u7684\u8d8b\u52bf\u662f minimax-m3 \u5728\u8be5\u65f6\u95f4\u7a97\u53e3\u5185\u4e0b\u964d\u4e86 55.1%\uff0c\u6700\u7ec8\u4e3a 4.59 \u8bcd\u5143/\u79d2\uff1bglm-5.2 \u6ce2\u52a8\u6027\u6700\u9ad8\uff0c\u5728 4.76 \u81f3 195.97 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 85.0%\u3002\n- \u65e0\u7f3a\u5931\u6570\u636e\u9650\u5236\uff1a\u5168\u90e8 8 \u4e2a\u6a21\u578b\u5747\u62a5\u544a\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u56e0\u6b64\u8be5\u56db\u5c0f\u65f6\u7a97\u53e3\u662f\u5b8c\u6574\u7684\u3002","total_duration_ns":6045724378,"translated_at":"2026-10-06T14:03:19.109019+00:00","translation_eval_count":185,"translation_model":"glm-5.3","translation_prompt_eval_count":336,"translation_total_duration_ns":11052058915,"translation_wall_seconds":11.2,"valid_point_count":96,"wall_seconds":6.236},{"coverage_pct":100.0,"eval_count":268,"generated_at":"2026-10-06T13:03:07.774538+00:00","generated_label":"Oct 06, 2026 \u00b7 13:03 UTC","id":1174,"model":"glm-5.3","period_end":"2026-10-06T13:03:01+00:00","period_label":"Oct 06, 2026 \u00b7 09:03 UTC to Oct 06, 2026 \u00b7 13:03 UTC","period_start":"2026-10-06T09:03:01+00:00","prompt_eval_count":2245,"summary":"- deepseek-v4.1-flash is the strongest model at 169.36 token/s average throughput (peak 230.12 token/s), while nemotron-3-ultra is the weakest at 17.71 token/s average (peak 25.64 token/s), a roughly tenfold gap between the two.\n- glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 83.5% and swings from 4.76 token/s at 11:00 to 211.18 token/s at 10:00, including a late dip to 7.84 token/s at 12:40 before recovering to 116.02 token/s.\n- No missing-data limitation exists in this window: all 8 models report 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour period.","summary_en":"- deepseek-v4.1-flash is the strongest model at 169.36 token/s average throughput (peak 230.12 token/s), while nemotron-3-ultra is the weakest at 17.71 token/s average (peak 25.64 token/s), a roughly tenfold gap between the two.\n- glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 83.5% and swings from 4.76 token/s at 11:00 to 211.18 token/s at 10:00, including a late dip to 7.84 token/s at 12:40 before recovering to 116.02 token/s.\n- No missing-data limitation exists in this window: all 8 models report 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour period.","summary_items":["deepseek-v4.1-flash is the strongest model at 169.36 token/s average throughput (peak 230.12 token/s), while nemotron-3-ultra is the weakest at 17.71 token/s average (peak 25.64 token/s), a roughly tenfold gap between the two.","glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 83.5% and swings from 4.76 token/s at 11:00 to 211.18 token/s at 10:00, including a late dip to 7.84 token/s at 12:40 before recovering to 116.02 token/s.","No missing-data limitation exists in this window: all 8 models report 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour period."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 169.36 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 230.12 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 17.71 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 25.64 \u8bcd\u5143/\u79d2\uff09\uff0c\u4e24\u8005\u4e4b\u95f4\u5b58\u5728\u7ea6\u5341\u500d\u7684\u5dee\u8ddd\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u5bf9\u8fd0\u884c\u5f71\u54cd\u6700\u4e3a\u663e\u8457\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 83.5%\uff0c\u5728 11:00 \u65f6\u4f4e\u81f3 4.76 \u8bcd\u5143/\u79d2\uff0c\u5728 10:00 \u65f6\u9ad8\u8fbe 211.18 \u8bcd\u5143/\u79d2\uff0c\u5176\u95f4\u5305\u62ec 12:40 \u65f6\u964d\u81f3 7.84 \u8bcd\u5143/\u79d2\u7684\u540e\u671f\u4f4e\u8c37\uff0c\u968f\u540e\u56de\u5347\u81f3 116.02 \u8bcd\u5143/\u79d2\u3002\n- \u8be5\u65f6\u95f4\u7a97\u53e3\u5185\u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8 8 \u4e2a\u6a21\u578b\u5747\u62a5\u544a\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u671f\u95f4\u5185\u5171\u4ea7\u751f 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":3392494988,"translated_at":"2026-10-06T13:03:07.774538+00:00","translation_eval_count":216,"translation_model":"glm-5.3","translation_prompt_eval_count":359,"translation_total_duration_ns":2730253207,"translation_wall_seconds":2.875,"valid_point_count":96,"wall_seconds":3.557},{"coverage_pct":100.0,"eval_count":464,"generated_at":"2026-10-06T12:03:07.940093+00:00","generated_label":"Oct 06, 2026 \u00b7 12:03 UTC","id":1173,"model":"glm-5.3","period_end":"2026-10-06T12:03:01+00:00","period_label":"Oct 06, 2026 \u00b7 08:03 UTC to Oct 06, 2026 \u00b7 12:03 UTC","period_start":"2026-10-06T08:03:01+00:00","prompt_eval_count":2245,"summary":"- deepseek-v4.1-flash is the strongest model at 174.26 token/s average throughput, peaking at 230.12 token/s; nemotron-3-ultra is the weakest at 19.79 token/s average, never exceeding 26.99 token/s.\n- glm-5.2 shows the most operationally significant volatility, with a 78.4% coefficient of variation and swings from 4.76 token/s at 11:00 to 211.18 token/s at 10:00; minimax-m3 also ranged between 12.2 and 74.93 token/s.\n- No missing-data limitation applies: all eight models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 174.26 token/s average throughput, peaking at 230.12 token/s; nemotron-3-ultra is the weakest at 19.79 token/s average, never exceeding 26.99 token/s.\n- glm-5.2 shows the most operationally significant volatility, with a 78.4% coefficient of variation and swings from 4.76 token/s at 11:00 to 211.18 token/s at 10:00; minimax-m3 also ranged between 12.2 and 74.93 token/s.\n- No missing-data limitation applies: all eight models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 174.26 token/s average throughput, peaking at 230.12 token/s; nemotron-3-ultra is the weakest at 19.79 token/s average, never exceeding 26.99 token/s.","glm-5.2 shows the most operationally significant volatility, with a 78.4% coefficient of variation and swings from 4.76 token/s at 11:00 to 211.18 token/s at 10:00; minimax-m3 also ranged between 12.2 and 74.93 token/s.","No missing-data limitation applies: all eight models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 174.26 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 230.12 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 19.79 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 26.99 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u5bf9\u8fd0\u884c\u5f71\u54cd\u6700\u4e3a\u663e\u8457\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 78.4%\uff0c\u5728 11:00 \u65f6\u4f4e\u81f3 4.76 \u8bcd\u5143/\u79d2\uff0c\u5728 10:00 \u65f6\u9ad8\u8fbe 211.18 \u8bcd\u5143/\u79d2\uff1bminimax-m3 \u7684\u6ce2\u52a8\u8303\u56f4\u4e5f\u5728 12.2 \u81f3 74.93 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62a5\u544a\u4e86 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u5171\u6709 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":3953388660,"translated_at":"2026-10-06T12:03:07.940093+00:00","translation_eval_count":195,"translation_model":"glm-5.3","translation_prompt_eval_count":339,"translation_total_duration_ns":2431345236,"translation_wall_seconds":2.58,"valid_point_count":96,"wall_seconds":4.122},{"coverage_pct":100.0,"eval_count":1006,"generated_at":"2026-10-06T11:03:07.446372+00:00","generated_label":"Oct 06, 2026 \u00b7 11:03 UTC","id":1172,"model":"glm-5.3","period_end":"2026-10-06T11:03:01+00:00","period_label":"Oct 06, 2026 \u00b7 07:03 UTC to Oct 06, 2026 \u00b7 11:03 UTC","period_start":"2026-10-06T07:03:01+00:00","prompt_eval_count":2247,"summary":"- deepseek-v4.1-flash is the strongest model at 172.39 token/s average throughput, peaking at 237.96 token/s; nemotron-3-ultra is the weakest at 22.34 token/s average, never exceeding 32.73 token/s.\n- glm-5.2 shows the most volatility (77.7% CV), spiking to 211.18 token/s at 10:00 UTC before collapsing to 4.76 token/s at 11:00 UTC; deepseek-v4.1-flash also dipped to 63.63 token/s at 09:40 UTC against its 172.39 token/s average.\n- No missing-data limitation exists: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 172.39 token/s average throughput, peaking at 237.96 token/s; nemotron-3-ultra is the weakest at 22.34 token/s average, never exceeding 32.73 token/s.\n- glm-5.2 shows the most volatility (77.7% CV), spiking to 211.18 token/s at 10:00 UTC before collapsing to 4.76 token/s at 11:00 UTC; deepseek-v4.1-flash also dipped to 63.63 token/s at 09:40 UTC against its 172.39 token/s average.\n- No missing-data limitation exists: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 172.39 token/s average throughput, peaking at 237.96 token/s; nemotron-3-ultra is the weakest at 22.34 token/s average, never exceeding 32.73 token/s.","glm-5.2 shows the most volatility (77.7% CV), spiking to 211.18 token/s at 10:00 UTC before collapsing to 4.76 token/s at 11:00 UTC; deepseek-v4.1-flash also dipped to 63.63 token/s at 09:40 UTC against its 172.39 token/s average.","No missing-data limitation exists: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 172.39 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 237.96 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 22.34 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 32.73 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u6ce2\u52a8\u6700\u5927\uff08CV \u4e3a 77.7%\uff09\uff0c\u5728 UTC 10:00 \u98d9\u5347\u81f3 211.18 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u5728 UTC 11:00 \u9aa4\u964d\u81f3 4.76 \u8bcd\u5143/\u79d2\uff1bdeepseek-v4.1-flash \u4e5f\u5728 UTC 09:40 \u8dcc\u81f3 63.63 \u8bcd\u5143/\u79d2\uff0c\u4f4e\u4e8e\u5176 172.39 \u8bcd\u5143/\u79d2\u7684\u5e73\u5747\u503c\u3002\n- \u4e0d\u5b58\u5728\u7f3a\u5931\u6570\u636e\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62a5\u544a 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\u300196 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5e76\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u5b9e\u73b0 100.0% \u7684\u8986\u76d6\u7387\u3002","total_duration_ns":4846069048,"translated_at":"2026-10-06T11:03:07.446372+00:00","translation_eval_count":211,"translation_model":"glm-5.3","translation_prompt_eval_count":350,"translation_total_duration_ns":999951468,"translation_wall_seconds":1.145,"valid_point_count":96,"wall_seconds":5.019},{"coverage_pct":100.0,"eval_count":799,"generated_at":"2026-10-06T10:03:10.921387+00:00","generated_label":"Oct 06, 2026 \u00b7 10:03 UTC","id":1171,"model":"glm-5.3","period_end":"2026-10-06T10:03:02+00:00","period_label":"Oct 06, 2026 \u00b7 06:03 UTC to Oct 06, 2026 \u00b7 10:03 UTC","period_start":"2026-10-06T06:03:02+00:00","prompt_eval_count":2248,"summary":"- deepseek-v4.1-flash is the strongest model at 167.71 token/s average throughput, peaking at 237.96 token/s, while nemotron-3-ultra is the weakest at 23.73 token/s average with a maximum of only 34.38 token/s.\n- glm-5.2 shows the most volatility, with a 61.5% coefficient of variation and swings from 15.88 to 211.18 token/s, including a spike to its maximum in the final 10:00 UTC sample; gemma4:31b declined 19.8% over the window.\n- No missing-data limitation exists: all eight models have 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 167.71 token/s average throughput, peaking at 237.96 token/s, while nemotron-3-ultra is the weakest at 23.73 token/s average with a maximum of only 34.38 token/s.\n- glm-5.2 shows the most volatility, with a 61.5% coefficient of variation and swings from 15.88 to 211.18 token/s, including a spike to its maximum in the final 10:00 UTC sample; gemma4:31b declined 19.8% over the window.\n- No missing-data limitation exists: all eight models have 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 167.71 token/s average throughput, peaking at 237.96 token/s, while nemotron-3-ultra is the weakest at 23.73 token/s average with a maximum of only 34.38 token/s.","glm-5.2 shows the most volatility, with a 61.5% coefficient of variation and swings from 15.88 to 211.18 token/s, including a spike to its maximum in the final 10:00 UTC sample; gemma4:31b declined 19.8% over the window.","No missing-data limitation exists: all eight models have 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 167.71 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 237.96 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4ec5 23.73 \u8bcd\u5143/\u79d2\uff0c\u6700\u5927\u503c\u4ec5\u4e3a 34.38 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u6ce2\u52a8\u6700\u5927\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 61.5%\uff0c\u6ce2\u52a8\u8303\u56f4\u4ece 15.88 \u5230 211.18 \u8bcd\u5143/\u79d2\uff0c\u5305\u62ec\u5728\u6700\u540e\u4e00\u4e2a 10:00 UTC \u91c7\u6837\u70b9\u98d9\u5347\u81f3\u6700\u5927\u503c\uff1bgemma4:31b \u5728\u8be5\u65f6\u95f4\u7a97\u53e3\u5185\u4e0b\u964d\u4e86 19.8%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u5171\u6709 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":6391032788,"translated_at":"2026-10-06T10:03:10.921387+00:00","translation_eval_count":184,"translation_model":"glm-5.3","translation_prompt_eval_count":341,"translation_total_duration_ns":2114336768,"translation_wall_seconds":2.259,"valid_point_count":96,"wall_seconds":6.557},{"coverage_pct":100.0,"eval_count":840,"generated_at":"2026-10-06T09:03:07.389069+00:00","generated_label":"Oct 06, 2026 \u00b7 09:03 UTC","id":1170,"model":"glm-5.3","period_end":"2026-10-06T09:03:02+00:00","period_label":"Oct 06, 2026 \u00b7 05:03 UTC to Oct 06, 2026 \u00b7 09:03 UTC","period_start":"2026-10-06T05:03:02+00:00","prompt_eval_count":2249,"summary":"- deepseek-v4.1-flash is the strongest model at 189.86 token/s average throughput (peak 237.96 token/s), while nemotron-3-ultra is the weakest at 25.15 token/s average (peak 34.38 token/s).\n- Volatility is the main operational concern: glm-5.2 shows 56.5% coefficient of variation with swings from 12.61 to 140.36 token/s, and minimax-m3 dropped to 5.15 token/s at 07:00; glm-5.3 is the only clear gainer, up 14.2% to 145.77 token/s at 09:00.\n- No missing-data limitation: all 8 models have 12 of 12 samples, 96 valid points, and 100.0% coverage over the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 189.86 token/s average throughput (peak 237.96 token/s), while nemotron-3-ultra is the weakest at 25.15 token/s average (peak 34.38 token/s).\n- Volatility is the main operational concern: glm-5.2 shows 56.5% coefficient of variation with swings from 12.61 to 140.36 token/s, and minimax-m3 dropped to 5.15 token/s at 07:00; glm-5.3 is the only clear gainer, up 14.2% to 145.77 token/s at 09:00.\n- No missing-data limitation: all 8 models have 12 of 12 samples, 96 valid points, and 100.0% coverage over the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 189.86 token/s average throughput (peak 237.96 token/s), while nemotron-3-ultra is the weakest at 25.15 token/s average (peak 34.38 token/s).","Volatility is the main operational concern: glm-5.2 shows 56.5% coefficient of variation with swings from 12.61 to 140.36 token/s, and minimax-m3 dropped to 5.15 token/s at 07:00; glm-5.3 is the only clear gainer, up 14.2% to 145.77 token/s at 09:00.","No missing-data limitation: all 8 models have 12 of 12 samples, 96 valid points, and 100.0% coverage over the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 189.86 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 237.96 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 25.15 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 34.38 \u8bcd\u5143/\u79d2\uff09\u3002\n- \u6ce2\u52a8\u6027\u662f\u4e3b\u8981\u7684\u8fd0\u8425\u95ee\u9898\uff1aglm-5.2 \u7684\u53d8\u5f02\u7cfb\u6570\u4e3a 56.5%\uff0c\u6ce2\u52a8\u8303\u56f4\u4ece 12.61 \u5230 140.36 \u8bcd\u5143/\u79d2\uff0cminimax-m3 \u5728 07:00 \u964d\u81f3 5.15 \u8bcd\u5143/\u79d2\uff1bglm-5.3 \u662f\u552f\u4e00\u660e\u663e\u63d0\u5347\u7684\u6a21\u578b\uff0c\u5728 09:00 \u63d0\u5347 14.2%\uff0c\u8fbe\u5230 145.77 \u8bcd\u5143/\u79d2\u3002\n- \u65e0\u7f3a\u5931\u6570\u636e\u9650\u5236\uff1a\u5168\u90e8 8 \u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":3972620977,"translated_at":"2026-10-06T09:03:07.389069+00:00","translation_eval_count":212,"translation_model":"glm-5.3","translation_prompt_eval_count":356,"translation_total_duration_ns":1007095949,"translation_wall_seconds":1.154,"valid_point_count":96,"wall_seconds":4.135},{"coverage_pct":100.0,"eval_count":328,"generated_at":"2026-10-06T08:03:07.347569+00:00","generated_label":"Oct 06, 2026 \u00b7 08:03 UTC","id":1169,"model":"glm-5.3","period_end":"2026-10-06T08:03:01+00:00","period_label":"Oct 06, 2026 \u00b7 04:03 UTC to Oct 06, 2026 \u00b7 08:03 UTC","period_start":"2026-10-06T04:03:01+00:00","prompt_eval_count":2249,"summary":"- deepseek-v4.1-flash is the strongest model with an average throughput of 184.76 token/s (peaking at 237.96 token/s), while nemotron-3-ultra is the weakest at 23.7 token/s average, never exceeding 34.38 token/s across the window.\n- The most significant volatility comes from glm-5.3, whose coefficient of variation is 42.3%, swinging from 171.53 token/s at 04:40 down to 6.45 token/s at 06:00; glm-5.2 shows the steepest trend, rising 95.8% from a 12.61 token/s low to a 140.36 token/s peak.\n- No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage over the four-hour period.","summary_en":"- deepseek-v4.1-flash is the strongest model with an average throughput of 184.76 token/s (peaking at 237.96 token/s), while nemotron-3-ultra is the weakest at 23.7 token/s average, never exceeding 34.38 token/s across the window.\n- The most significant volatility comes from glm-5.3, whose coefficient of variation is 42.3%, swinging from 171.53 token/s at 04:40 down to 6.45 token/s at 06:00; glm-5.2 shows the steepest trend, rising 95.8% from a 12.61 token/s low to a 140.36 token/s peak.\n- No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage over the four-hour period.","summary_items":["deepseek-v4.1-flash is the strongest model with an average throughput of 184.76 token/s (peaking at 237.96 token/s), while nemotron-3-ultra is the weakest at 23.7 token/s average, never exceeding 34.38 token/s across the window.","The most significant volatility comes from glm-5.3, whose coefficient of variation is 42.3%, swinging from 171.53 token/s at 04:40 down to 6.45 token/s at 06:00; glm-5.2 shows the steepest trend, rising 95.8% from a 12.61 token/s low to a 140.36 token/s peak.","No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage over the four-hour period."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 184.76 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c\u8fbe 237.96 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4ec5 23.7 \u8bcd\u5143/\u79d2\uff0c\u5728\u6574\u4e2a\u65f6\u95f4\u7a97\u53e3\u5185\u4ece\u672a\u8d85\u8fc7 34.38 \u8bcd\u5143/\u79d2\u3002\n- \u6ce2\u52a8\u6700\u663e\u8457\u7684\u662f glm-5.3\uff0c\u5176\u53d8\u5f02\u7cfb\u6570\u4e3a 42.3%\uff0c\u4ece 04:40 \u7684 171.53 \u8bcd\u5143/\u79d2\u9aa4\u964d\u81f3 06:00 \u7684 6.45 \u8bcd\u5143/\u79d2\uff1bglm-5.2 \u8d8b\u52bf\u6700\u4e3a\u9661\u5ced\uff0c\u4ece 12.61 \u8bcd\u5143/\u79d2\u7684\u4f4e\u70b9\u5347\u81f3 140.36 \u8bcd\u5143/\u79d2\u7684\u5cf0\u503c\uff0c\u6da8\u5e45\u8fbe 95.8%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u4ea4\u4ed8\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u5185\u5171\u6709 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":3627462997,"translated_at":"2026-10-06T08:03:07.347569+00:00","translation_eval_count":213,"translation_model":"glm-5.3","translation_prompt_eval_count":364,"translation_total_duration_ns":1980749774,"translation_wall_seconds":2.127,"valid_point_count":96,"wall_seconds":3.789},{"coverage_pct":100.0,"eval_count":284,"generated_at":"2026-10-06T07:03:07.976370+00:00","generated_label":"Oct 06, 2026 \u00b7 07:03 UTC","id":1168,"model":"glm-5.3","period_end":"2026-10-06T07:03:01+00:00","period_label":"Oct 06, 2026 \u00b7 03:03 UTC to Oct 06, 2026 \u00b7 07:03 UTC","period_start":"2026-10-06T03:03:01+00:00","prompt_eval_count":2252,"summary":"- deepseek-v4.1-flash is the strongest model at 197.65 token/s average throughput (peaking at 231.48 token/s), while nemotron-3-ultra is the weakest at 24.70 token/s average, never exceeding 36.63 token/s.\n- glm-5.2 shows the most volatility, with a 73.3% coefficient of variation and swings from 12.61 to 218.07 token/s; minimax-m3 also fell sharply to 5.15 token/s at 07:00, its lowest reading.\n- No missing-data limitation exists: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 197.65 token/s average throughput (peaking at 231.48 token/s), while nemotron-3-ultra is the weakest at 24.70 token/s average, never exceeding 36.63 token/s.\n- glm-5.2 shows the most volatility, with a 73.3% coefficient of variation and swings from 12.61 to 218.07 token/s; minimax-m3 also fell sharply to 5.15 token/s at 07:00, its lowest reading.\n- No missing-data limitation exists: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 197.65 token/s average throughput (peaking at 231.48 token/s), while nemotron-3-ultra is the weakest at 24.70 token/s average, never exceeding 36.63 token/s.","glm-5.2 shows the most volatility, with a 73.3% coefficient of variation and swings from 12.61 to 218.07 token/s; minimax-m3 also fell sharply to 5.15 token/s at 07:00, its lowest reading.","No missing-data limitation exists: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u8fbe 197.65 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 231.48 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4ec5 24.70 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 36.63 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u6ce2\u52a8\u6700\u5927\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 73.3%\uff0c\u5728 12.61 \u81f3 218.07 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u5927\u5e45\u6ce2\u52a8\uff1bminimax-m3 \u4e5f\u5728 07:00 \u9aa4\u964d\u81f3 5.15 \u8bcd\u5143/\u79d2\uff0c\u4e3a\u5176\u6700\u4f4e\u8bfb\u6570\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u95ee\u9898\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0cvalid_point_count \u4e3a 96\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":2731912058,"translated_at":"2026-10-06T07:03:07.976370+00:00","translation_eval_count":179,"translation_model":"glm-5.3","translation_prompt_eval_count":336,"translation_total_duration_ns":2985144010,"translation_wall_seconds":3.133,"valid_point_count":96,"wall_seconds":2.891},{"coverage_pct":100.0,"eval_count":783,"generated_at":"2026-10-06T06:03:05.958600+00:00","generated_label":"Oct 06, 2026 \u00b7 06:03 UTC","id":1167,"model":"glm-5.3","period_end":"2026-10-06T06:03:01+00:00","period_label":"Oct 06, 2026 \u00b7 02:03 UTC to Oct 06, 2026 \u00b7 06:03 UTC","period_start":"2026-10-06T02:03:01+00:00","prompt_eval_count":2251,"summary":"- deepseek-v4.1-flash is the strongest model at 189.69 token/s average throughput (peaking at 229.56 token/s), while nemotron-3-ultra is the weakest at 24.74 token/s average, roughly eight times slower.\n- glm-5.2 shows the most volatility, with a coefficient of variation of 89.5% and swings from 218.07 token/s at 04:00 down to 12.61 token/s at 05:20; glm-5.3 also dropped sharply to 6.45 token/s in the final 06:00 observation.\n- No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 189.69 token/s average throughput (peaking at 229.56 token/s), while nemotron-3-ultra is the weakest at 24.74 token/s average, roughly eight times slower.\n- glm-5.2 shows the most volatility, with a coefficient of variation of 89.5% and swings from 218.07 token/s at 04:00 down to 12.61 token/s at 05:20; glm-5.3 also dropped sharply to 6.45 token/s in the final 06:00 observation.\n- No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 189.69 token/s average throughput (peaking at 229.56 token/s), while nemotron-3-ultra is the weakest at 24.74 token/s average, roughly eight times slower.","glm-5.2 shows the most volatility, with a coefficient of variation of 89.5% and swings from 218.07 token/s at 04:00 down to 12.61 token/s at 05:20; glm-5.3 also dropped sharply to 6.45 token/s in the final 06:00 observation.","No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 189.69 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c\u8fbe 229.56 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 24.74 \u8bcd\u5143/\u79d2\uff0c\u5927\u7ea6\u6162\u516b\u500d\u3002\n- glm-5.2 \u6ce2\u52a8\u6700\u5927\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 89.5%\uff0c\u4ece 04:00 \u7684 218.07 \u8bcd\u5143/\u79d2\u9aa4\u964d\u81f3 05:20 \u7684 12.61 \u8bcd\u5143/\u79d2\uff1bglm-5.3 \u5728\u6700\u540e 06:00 \u7684\u89c2\u6d4b\u4e2d\u4e5f\u6025\u5267\u4e0b\u964d\u81f3 6.45 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0cvalid_point_count \u4e3a 96\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":3490628363,"translated_at":"2026-10-06T06:03:05.958600+00:00","translation_eval_count":192,"translation_model":"glm-5.3","translation_prompt_eval_count":344,"translation_total_duration_ns":956907381,"translation_wall_seconds":1.098,"valid_point_count":96,"wall_seconds":3.648},{"coverage_pct":100.0,"eval_count":524,"generated_at":"2026-10-06T05:03:05.743929+00:00","generated_label":"Oct 06, 2026 \u00b7 05:03 UTC","id":1166,"model":"glm-5.3","period_end":"2026-10-06T05:03:01+00:00","period_label":"Oct 06, 2026 \u00b7 01:03 UTC to Oct 06, 2026 \u00b7 05:03 UTC","period_start":"2026-10-06T01:03:01+00:00","prompt_eval_count":2249,"summary":"- deepseek-v4.1-flash is the strongest model at 177.60 token/s average throughput (range 122.23 to 229.51 token/s), while nemotron-3-ultra is the weakest at 26.83 token/s average (range 10.18 to 36.63 token/s).\n- glm-5.2 shows the most operationally significant volatility, with a 78.9% coefficient of variation and swings from 20.74 token/s at 03:00 to 218.07 token/s at 04:00; deepseek-v4.1-flash is the steadiest at 18.4%.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset records 96 valid points with 100.0% coverage, though it covers only this four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 177.60 token/s average throughput (range 122.23 to 229.51 token/s), while nemotron-3-ultra is the weakest at 26.83 token/s average (range 10.18 to 36.63 token/s).\n- glm-5.2 shows the most operationally significant volatility, with a 78.9% coefficient of variation and swings from 20.74 token/s at 03:00 to 218.07 token/s at 04:00; deepseek-v4.1-flash is the steadiest at 18.4%.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset records 96 valid points with 100.0% coverage, though it covers only this four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 177.60 token/s average throughput (range 122.23 to 229.51 token/s), while nemotron-3-ultra is the weakest at 26.83 token/s average (range 10.18 to 36.63 token/s).","glm-5.2 shows the most operationally significant volatility, with a 78.9% coefficient of variation and swings from 20.74 token/s at 03:00 to 218.07 token/s at 04:00; deepseek-v4.1-flash is the steadiest at 18.4%.","No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset records 96 valid points with 100.0% coverage, though it covers only this four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 177.60 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 122.23 \u81f3 229.51 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 26.83 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 10.18 \u81f3 36.63 \u8bcd\u5143/\u79d2\uff09\u3002\n- glm-5.2 \u7684\u6ce2\u52a8\u6027\u5bf9\u8fd0\u8425\u5f71\u54cd\u6700\u5927\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 78.9%\uff0c\u4ece 03:00 \u7684 20.74 \u8bcd\u5143/\u79d2\u6ce2\u52a8\u81f3 04:00 \u7684 218.07 \u8bcd\u5143/\u79d2\uff1bdeepseek-v4.1-flash \u6700\u4e3a\u7a33\u5b9a\uff0c\u4e3a 18.4%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u6570\u636e\u96c6\u8bb0\u5f55\u4e86 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u4f46\u5176\u4ec5\u8986\u76d6\u8fd9\u56db\u4e2a\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u3002","total_duration_ns":2451567433,"translated_at":"2026-10-06T05:03:05.743929+00:00","translation_eval_count":200,"translation_model":"glm-5.3","translation_prompt_eval_count":355,"translation_total_duration_ns":1178581607,"translation_wall_seconds":1.326,"valid_point_count":96,"wall_seconds":2.619},{"coverage_pct":100.0,"eval_count":794,"generated_at":"2026-10-06T04:03:06.300167+00:00","generated_label":"Oct 06, 2026 \u00b7 04:03 UTC","id":1165,"model":"glm-5.3","period_end":"2026-10-06T04:03:01+00:00","period_label":"Oct 06, 2026 \u00b7 00:03 UTC to Oct 06, 2026 \u00b7 04:03 UTC","period_start":"2026-10-06T00:03:01+00:00","prompt_eval_count":2251,"summary":"- deepseek-v4.1-flash is the strongest model at 186.86 token/s average throughput, peaking at 229.51 token/s, while nemotron-3-ultra is the weakest at 29.10 token/s average, never exceeding 36.63 token/s.\n- glm-5.2 shows the most operationally significant volatility, with a 73.9% coefficient of variation and swings from 20.74 to 218.07 token/s; it also declined 26.3% over the window, and glm-5.3 fell 35.4%.\n- No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 186.86 token/s average throughput, peaking at 229.51 token/s, while nemotron-3-ultra is the weakest at 29.10 token/s average, never exceeding 36.63 token/s.\n- glm-5.2 shows the most operationally significant volatility, with a 73.9% coefficient of variation and swings from 20.74 to 218.07 token/s; it also declined 26.3% over the window, and glm-5.3 fell 35.4%.\n- No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 186.86 token/s average throughput, peaking at 229.51 token/s, while nemotron-3-ultra is the weakest at 29.10 token/s average, never exceeding 36.63 token/s.","glm-5.2 shows the most operationally significant volatility, with a 73.9% coefficient of variation and swings from 20.74 to 218.07 token/s; it also declined 26.3% over the window, and glm-5.3 fell 35.4%.","No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 186.86 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 229.51 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 29.10 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 36.63 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u5bf9\u8fd0\u8425\u5f71\u54cd\u6700\u5927\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 73.9%\uff0c\u6ce2\u52a8\u8303\u56f4\u4ece 20.74 \u5230 218.07 \u8bcd\u5143/\u79d2\uff1b\u5b83\u5728\u8be5\u65f6\u95f4\u7a97\u53e3\u5185\u8fd8\u4e0b\u964d\u4e86 26.3%\uff0cglm-5.3 \u4e0b\u964d\u4e86 35.4%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u8bb0\u5f55\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u4ea7\u751f 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u8986\u76d6\u7387\u8fbe\u5230 100.0%\u3002","total_duration_ns":3178511092,"translated_at":"2026-10-06T04:03:06.300167+00:00","translation_eval_count":193,"translation_model":"glm-5.3","translation_prompt_eval_count":336,"translation_total_duration_ns":938941186,"translation_wall_seconds":1.09,"valid_point_count":96,"wall_seconds":3.348},{"coverage_pct":100.0,"eval_count":768,"generated_at":"2026-10-06T03:03:06.651528+00:00","generated_label":"Oct 06, 2026 \u00b7 03:03 UTC","id":1164,"model":"glm-5.3","period_end":"2026-10-06T03:03:01+00:00","period_label":"Oct 05, 2026 \u00b7 23:03 UTC to Oct 06, 2026 \u00b7 03:03 UTC","period_start":"2026-10-05T23:03:01+00:00","prompt_eval_count":2244,"summary":"- deepseek-v4.1-flash is the strongest model at 175.25 token/s average throughput (peak 211.09 token/s), while nemotron-3-ultra is the weakest at 31.75 token/s average, peaking at only 58.87 token/s.\n- glm-5.3 shows the sharpest decline, trending -37.8% from roughly 180 token/s early in the window to a 53.18 token/s low at 02:20; glm-5.2 is the most volatile, with a 75.7% coefficient of variation spanning 11.08 to 198.51 token/s.\n- No missing-data limitation exists: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 175.25 token/s average throughput (peak 211.09 token/s), while nemotron-3-ultra is the weakest at 31.75 token/s average, peaking at only 58.87 token/s.\n- glm-5.3 shows the sharpest decline, trending -37.8% from roughly 180 token/s early in the window to a 53.18 token/s low at 02:20; glm-5.2 is the most volatile, with a 75.7% coefficient of variation spanning 11.08 to 198.51 token/s.\n- No missing-data limitation exists: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 175.25 token/s average throughput (peak 211.09 token/s), while nemotron-3-ultra is the weakest at 31.75 token/s average, peaking at only 58.87 token/s.","glm-5.3 shows the sharpest decline, trending -37.8% from roughly 180 token/s early in the window to a 53.18 token/s low at 02:20; glm-5.2 is the most volatile, with a 75.7% coefficient of variation spanning 11.08 to 198.51 token/s.","No missing-data limitation exists: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 175.25 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 211.09 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4ec5\u4e3a 31.75 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u53ea\u6709 58.87 \u8bcd\u5143/\u79d2\u3002\n- glm-5.3 \u4e0b\u964d\u6700\u4e3a\u6025\u5267\uff0c\u4ece\u7a97\u53e3\u521d\u671f\u7ea6 180 \u8bcd\u5143/\u79d2\u7684\u8d70\u52bf\u4e0b\u6ed1 -37.8%\uff0c\u5728 02:20 \u964d\u81f3 53.18 \u8bcd\u5143/\u79d2\u7684\u4f4e\u70b9\uff1bglm-5.2 \u6ce2\u52a8\u6027\u6700\u5927\uff0c\u53d8\u5f02\u7cfb\u6570\u8fbe 75.7%\uff0c\u6ce2\u52a8\u8303\u56f4\u4ecb\u4e8e 11.08 \u81f3 198.51 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u3002\n- \u4e0d\u5b58\u5728\u7f3a\u5931\u6570\u636e\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\u300196 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u8986\u76d6\u7387\u8fbe\u5230 100.0%\u3002","total_duration_ns":3789945749,"translated_at":"2026-10-06T03:03:06.651528+00:00","translation_eval_count":200,"translation_model":"glm-5.3","translation_prompt_eval_count":346,"translation_total_duration_ns":936563928,"translation_wall_seconds":1.087,"valid_point_count":96,"wall_seconds":3.956},{"coverage_pct":100.0,"eval_count":744,"generated_at":"2026-10-06T02:03:05.731829+00:00","generated_label":"Oct 06, 2026 \u00b7 02:03 UTC","id":1163,"model":"glm-5.3","period_end":"2026-10-06T02:03:01+00:00","period_label":"Oct 05, 2026 \u00b7 22:03 UTC to Oct 06, 2026 \u00b7 02:03 UTC","period_start":"2026-10-05T22:03:01+00:00","prompt_eval_count":2246,"summary":"- deepseek-v4.1-flash is the strongest model at 182.38 token/s average throughput (peaking at 213.39 token/s), while nemotron-3-ultra is the weakest at 32.17 token/s average (minimum 11.42 token/s).\n- glm-5.2 shows the most volatility, with a 69.3% coefficient of variation and swings from 11.08 to 198.51 token/s, including a 41.0% upward trend; by contrast deepseek-v4.1-flash stays stable with only 13.5% variation.\n- No missing-data limitation exists: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 182.38 token/s average throughput (peaking at 213.39 token/s), while nemotron-3-ultra is the weakest at 32.17 token/s average (minimum 11.42 token/s).\n- glm-5.2 shows the most volatility, with a 69.3% coefficient of variation and swings from 11.08 to 198.51 token/s, including a 41.0% upward trend; by contrast deepseek-v4.1-flash stays stable with only 13.5% variation.\n- No missing-data limitation exists: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 182.38 token/s average throughput (peaking at 213.39 token/s), while nemotron-3-ultra is the weakest at 32.17 token/s average (minimum 11.42 token/s).","glm-5.2 shows the most volatility, with a 69.3% coefficient of variation and swings from 11.08 to 198.51 token/s, including a 41.0% upward trend; by contrast deepseek-v4.1-flash stays stable with only 13.5% variation.","No missing-data limitation exists: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 182.38 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c\u8fbe 213.39 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 32.17 \u8bcd\u5143/\u79d2\uff08\u6700\u4f4e 11.42 \u8bcd\u5143/\u79d2\uff09\u3002\n- glm-5.2 \u6ce2\u52a8\u6700\u5927\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 69.3%\uff0c\u6ce2\u52a8\u8303\u56f4\u4ece 11.08 \u5230 198.51 \u8bcd\u5143/\u79d2\uff0c\u5305\u62ec 41.0% \u7684\u4e0a\u5347\u8d8b\u52bf\uff1b\u76f8\u6bd4\u4e4b\u4e0b\uff0cdeepseek-v4.1-flash \u4fdd\u6301\u7a33\u5b9a\uff0c\u53d8\u5f02\u4ec5\u4e3a 13.5%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u6240\u6709 8 \u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0cvalid_point_count \u4e3a 96\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":3114120909,"translated_at":"2026-10-06T02:03:05.731829+00:00","translation_eval_count":183,"translation_model":"glm-5.3","translation_prompt_eval_count":339,"translation_total_duration_ns":918827299,"translation_wall_seconds":1.234,"valid_point_count":96,"wall_seconds":3.289},{"coverage_pct":100.0,"eval_count":282,"generated_at":"2026-10-06T01:03:04.271273+00:00","generated_label":"Oct 06, 2026 \u00b7 01:03 UTC","id":1162,"model":"glm-5.3","period_end":"2026-10-06T01:03:01+00:00","period_label":"Oct 05, 2026 \u00b7 21:03 UTC to Oct 06, 2026 \u00b7 01:03 UTC","period_start":"2026-10-05T21:03:01+00:00","prompt_eval_count":2246,"summary":"- deepseek-v4.1-flash is the strongest model at 184.68 token/s average throughput (peak 213.39 token/s), while nemotron-3-ultra is the weakest at 29.10 token/s average, roughly six times slower.\n- glm-5.3 shows the steadiest improvement, rising 33.4% to a 153.15 token/s average with only 19.5% coefficient of variation, whereas glm-5.2 is the most volatile at 86.1% coefficient of variation, swinging between 11.08 and 198.51 token/s.\n- No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 184.68 token/s average throughput (peak 213.39 token/s), while nemotron-3-ultra is the weakest at 29.10 token/s average, roughly six times slower.\n- glm-5.3 shows the steadiest improvement, rising 33.4% to a 153.15 token/s average with only 19.5% coefficient of variation, whereas glm-5.2 is the most volatile at 86.1% coefficient of variation, swinging between 11.08 and 198.51 token/s.\n- No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 184.68 token/s average throughput (peak 213.39 token/s), while nemotron-3-ultra is the weakest at 29.10 token/s average, roughly six times slower.","glm-5.3 shows the steadiest improvement, rising 33.4% to a 153.15 token/s average with only 19.5% coefficient of variation, whereas glm-5.2 is the most volatile at 86.1% coefficient of variation, swinging between 11.08 and 198.51 token/s.","No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 184.68 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 213.39 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 29.10 \u8bcd\u5143/\u79d2\uff0c\u5927\u7ea6\u6162\u516d\u500d\u3002\n- glm-5.3 \u8868\u73b0\u51fa\u6700\u7a33\u5b9a\u7684\u63d0\u5347\uff0c\u4e0a\u5347 33.4% \u81f3\u5e73\u5747 153.15 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4ec5\u4e3a 19.5%\uff1b\u800c glm-5.2 \u6ce2\u52a8\u6700\u5927\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 86.1%\uff0c\u5728 11.08 \u81f3 198.51 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u4ea4\u4ed8\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u5171\u6709 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":1471175526,"translated_at":"2026-10-06T01:03:04.271273+00:00","translation_eval_count":201,"translation_model":"glm-5.3","translation_prompt_eval_count":338,"translation_total_duration_ns":936932075,"translation_wall_seconds":1.088,"valid_point_count":96,"wall_seconds":1.635},{"coverage_pct":100.0,"eval_count":291,"generated_at":"2026-10-06T00:03:04.126787+00:00","generated_label":"Oct 06, 2026 \u00b7 00:03 UTC","id":1161,"model":"glm-5.3","period_end":"2026-10-06T00:03:01+00:00","period_label":"Oct 05, 2026 \u00b7 20:03 UTC to Oct 06, 2026 \u00b7 00:03 UTC","period_start":"2026-10-05T20:03:01+00:00","prompt_eval_count":2246,"summary":"- deepseek-v4.1-flash is the strongest model with an average throughput of 180.61 token/s (p95 215.23 token/s, max 217.48 token/s), while nemotron-3-ultra is the weakest at 27.23 token/s average, peaking at only 58.87 token/s.\n- glm-5.2 shows the most volatility, with a coefficient of variation of 77.7% and swings from 11.08 token/s at 00:00 to a 191.19 token/s spike at 22:20; glm-5.3-flash also swung between 34.83 and 185.87 token/s within the window.\n- No missing-data limitation applies: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0% across the four-hour period, so the dataset is complete.","summary_en":"- deepseek-v4.1-flash is the strongest model with an average throughput of 180.61 token/s (p95 215.23 token/s, max 217.48 token/s), while nemotron-3-ultra is the weakest at 27.23 token/s average, peaking at only 58.87 token/s.\n- glm-5.2 shows the most volatility, with a coefficient of variation of 77.7% and swings from 11.08 token/s at 00:00 to a 191.19 token/s spike at 22:20; glm-5.3-flash also swung between 34.83 and 185.87 token/s within the window.\n- No missing-data limitation applies: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0% across the four-hour period, so the dataset is complete.","summary_items":["deepseek-v4.1-flash is the strongest model with an average throughput of 180.61 token/s (p95 215.23 token/s, max 217.48 token/s), while nemotron-3-ultra is the weakest at 27.23 token/s average, peaking at only 58.87 token/s.","glm-5.2 shows the most volatility, with a coefficient of variation of 77.7% and swings from 11.08 token/s at 00:00 to a 191.19 token/s spike at 22:20; glm-5.3-flash also swung between 34.83 and 185.87 token/s within the window.","No missing-data limitation applies: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0% across the four-hour period, so the dataset is complete."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 180.61 \u8bcd\u5143/\u79d2\uff08p95 215.23 \u8bcd\u5143/\u79d2\uff0c\u6700\u5927 217.48 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 27.23 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u4ec5\u4e3a 58.87 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u6ce2\u52a8\u6700\u5927\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 77.7%\uff0c\u4ece 00:00 \u7684 11.08 \u8bcd\u5143/\u79d2\u6ce2\u52a8\u5230 22:20 \u7684 191.19 \u8bcd\u5143/\u79d2\u5cf0\u503c\uff1bglm-5.3-flash \u5728\u8be5\u65f6\u95f4\u7a97\u53e3\u5185\u4e5f\u5728 34.83 \u81f3 185.87 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0cvalid_point_count \u4e3a 96\uff0c\u5728\u56db\u5c0f\u65f6\u671f\u95f4\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u56e0\u6b64\u6570\u636e\u96c6\u662f\u5b8c\u6574\u7684\u3002","total_duration_ns":1406785596,"translated_at":"2026-10-06T00:03:04.126787+00:00","translation_eval_count":216,"translation_model":"glm-5.3","translation_prompt_eval_count":367,"translation_total_duration_ns":847791509,"translation_wall_seconds":1.03,"valid_point_count":96,"wall_seconds":1.75},{"coverage_pct":100.0,"eval_count":963,"generated_at":"2026-10-05T23:03:08.117710+00:00","generated_label":"Oct 05, 2026 \u00b7 23:03 UTC","id":1160,"model":"glm-5.3","period_end":"2026-10-05T23:03:02+00:00","period_label":"Oct 05, 2026 \u00b7 19:03 UTC to Oct 05, 2026 \u00b7 23:03 UTC","period_start":"2026-10-05T19:03:02+00:00","prompt_eval_count":2249,"summary":"- deepseek-v4.1-flash is the strongest model at 183.38 token/s average throughput (peak 227.33 token/s), while nemotron-3-ultra is the weakest at 21.14 token/s average, roughly one-ninth of the leader.\n- glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 73.7 percent, swinging from 19.99 token/s at 21:40 to 191.19 token/s at 22:20; gemma4:31b also dipped to 50.18 token/s around 20:40.\n- No missing-data limitation applies: all eight models have 12 of 12 samples and 100.0 percent coverage, with 96 valid points overall, so the four-hour window is fully represented.","summary_en":"- deepseek-v4.1-flash is the strongest model at 183.38 token/s average throughput (peak 227.33 token/s), while nemotron-3-ultra is the weakest at 21.14 token/s average, roughly one-ninth of the leader.\n- glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 73.7 percent, swinging from 19.99 token/s at 21:40 to 191.19 token/s at 22:20; gemma4:31b also dipped to 50.18 token/s around 20:40.\n- No missing-data limitation applies: all eight models have 12 of 12 samples and 100.0 percent coverage, with 96 valid points overall, so the four-hour window is fully represented.","summary_items":["deepseek-v4.1-flash is the strongest model at 183.38 token/s average throughput (peak 227.33 token/s), while nemotron-3-ultra is the weakest at 21.14 token/s average, roughly one-ninth of the leader.","glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 73.7 percent, swinging from 19.99 token/s at 21:40 to 191.19 token/s at 22:20; gemma4:31b also dipped to 50.18 token/s around 20:40.","No missing-data limitation applies: all eight models have 12 of 12 samples and 100.0 percent coverage, with 96 valid points overall, so the four-hour window is fully represented."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 183.38 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 227.33 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 21.14 \u8bcd\u5143/\u79d2\uff0c\u7ea6\u4e3a\u9886\u5148\u8005\u7684\u4e5d\u5206\u4e4b\u4e00\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u5bf9\u8fd0\u8425\u5f71\u54cd\u6700\u5927\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 73.7%\uff0c\u4ece 21:40 \u7684 19.99 \u8bcd\u5143/\u79d2\u6ce2\u52a8\u5230 22:20 \u7684 191.19 \u8bcd\u5143/\u79d2\uff1bgemma4:31b \u4e5f\u5728 20:40 \u524d\u540e\u8dcc\u81f3 50.18 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u603b\u8ba1 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u56e0\u6b64\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5f97\u5230\u4e86\u5b8c\u6574\u5448\u73b0\u3002","total_duration_ns":4534570929,"translated_at":"2026-10-05T23:03:08.117710+00:00","translation_eval_count":214,"translation_model":"glm-5.3","translation_prompt_eval_count":347,"translation_total_duration_ns":1075359756,"translation_wall_seconds":1.22,"valid_point_count":96,"wall_seconds":4.875},{"coverage_pct":100.0,"eval_count":828,"generated_at":"2026-10-05T22:03:07.658969+00:00","generated_label":"Oct 05, 2026 \u00b7 22:03 UTC","id":1159,"model":"glm-5.3","period_end":"2026-10-05T22:03:01+00:00","period_label":"Oct 05, 2026 \u00b7 18:03 UTC to Oct 05, 2026 \u00b7 22:03 UTC","period_start":"2026-10-05T18:03:01+00:00","prompt_eval_count":2247,"summary":"- deepseek-v4.1-flash is the strongest model at 180.13 token/s average throughput (range 129.09\u2013227.33 token/s), while nemotron-3-ultra is the weakest at 21.11 token/s average, never exceeding 41.32 token/s.\n- glm-5.3 shows the steadiest gain, trending +19.0% to a 123.86 token/s average with the lowest volatility (cv 22.4%); glm-5.2 is the most volatile (cv 57.4%), swinging between 19.99 and 125.78 token/s, and minimax-m3 dipped to 6.95 token/s at 19:20.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, with 96 valid points and 100.0% coverage, so the four-hour window is fully represented.","summary_en":"- deepseek-v4.1-flash is the strongest model at 180.13 token/s average throughput (range 129.09\u2013227.33 token/s), while nemotron-3-ultra is the weakest at 21.11 token/s average, never exceeding 41.32 token/s.\n- glm-5.3 shows the steadiest gain, trending +19.0% to a 123.86 token/s average with the lowest volatility (cv 22.4%); glm-5.2 is the most volatile (cv 57.4%), swinging between 19.99 and 125.78 token/s, and minimax-m3 dipped to 6.95 token/s at 19:20.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, with 96 valid points and 100.0% coverage, so the four-hour window is fully represented.","summary_items":["deepseek-v4.1-flash is the strongest model at 180.13 token/s average throughput (range 129.09\u2013227.33 token/s), while nemotron-3-ultra is the weakest at 21.11 token/s average, never exceeding 41.32 token/s.","glm-5.3 shows the steadiest gain, trending +19.0% to a 123.86 token/s average with the lowest volatility (cv 22.4%); glm-5.2 is the most volatile (cv 57.4%), swinging between 19.99 and 125.78 token/s, and minimax-m3 dipped to 6.95 token/s at 19:20.","No missing-data limitation applies: all eight models have 12 of 12 samples, with 96 valid points and 100.0% coverage, so the four-hour window is fully represented."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 180.13 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 129.09\u2013227.33 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 21.11 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 41.32 \u8bcd\u5143/\u79d2\u3002\n- glm-5.3 \u589e\u957f\u6700\u4e3a\u7a33\u5b9a\uff0c\u5448 +19.0% \u7684\u4e0a\u5347\u8d8b\u52bf\uff0c\u5e73\u5747\u8fbe\u5230 123.86 \u8bcd\u5143/\u79d2\uff0c\u4e14\u6ce2\u52a8\u6027\u6700\u4f4e\uff08\u53d8\u5f02\u7cfb\u6570 22.4%\uff09\uff1bglm-5.2 \u6ce2\u52a8\u6027\u6700\u5927\uff08\u53d8\u5f02\u7cfb\u6570 57.4%\uff09\uff0c\u5728 19.99 \u81f3 125.78 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6446\u52a8\uff0cminimax-m3 \u5728 19:20 \u964d\u81f3 6.95 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u56e0\u6b64\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5f97\u5230\u4e86\u5b8c\u6574\u5448\u73b0\u3002","total_duration_ns":4328721566,"translated_at":"2026-10-05T22:03:07.658969+00:00","translation_eval_count":234,"translation_model":"glm-5.3","translation_prompt_eval_count":366,"translation_total_duration_ns":1217856590,"translation_wall_seconds":1.371,"valid_point_count":96,"wall_seconds":4.677},{"coverage_pct":100.0,"eval_count":533,"generated_at":"2026-10-05T21:03:04.948016+00:00","generated_label":"Oct 05, 2026 \u00b7 21:03 UTC","id":1158,"model":"glm-5.3","period_end":"2026-10-05T21:03:01+00:00","period_label":"Oct 05, 2026 \u00b7 17:03 UTC to Oct 05, 2026 \u00b7 21:03 UTC","period_start":"2026-10-05T17:03:01+00:00","prompt_eval_count":2247,"summary":"- deepseek-v4.1-flash is the strongest model at 163.48 token/s average throughput, peaking at 227.33 token/s, while nemotron-3-ultra is the weakest at 20.78 token/s average and never exceeding 38.18 token/s.\n- The most operationally significant trend is glm-5.3's 78.7% rise, from 49.21 token/s at 18:00 to 143.23 token/s at 20:00; glm-5.2 shows the greatest volatility with a 48.3% coefficient of variation, swinging between 31.0 and 125.78 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, so figures represent the complete four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 163.48 token/s average throughput, peaking at 227.33 token/s, while nemotron-3-ultra is the weakest at 20.78 token/s average and never exceeding 38.18 token/s.\n- The most operationally significant trend is glm-5.3's 78.7% rise, from 49.21 token/s at 18:00 to 143.23 token/s at 20:00; glm-5.2 shows the greatest volatility with a 48.3% coefficient of variation, swinging between 31.0 and 125.78 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, so figures represent the complete four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 163.48 token/s average throughput, peaking at 227.33 token/s, while nemotron-3-ultra is the weakest at 20.78 token/s average and never exceeding 38.18 token/s.","The most operationally significant trend is glm-5.3's 78.7% rise, from 49.21 token/s at 18:00 to 143.23 token/s at 20:00; glm-5.2 shows the greatest volatility with a 48.3% coefficient of variation, swinging between 31.0 and 125.78 token/s.","No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, so figures represent the complete four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 163.48 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 227.33 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 20.78 \u8bcd\u5143/\u79d2\uff0c\u4e14\u4ece\u672a\u8d85\u8fc7 38.18 \u8bcd\u5143/\u79d2\u3002\n- \u8fd0\u8425\u5c42\u9762\u6700\u663e\u8457\u7684\u8d8b\u52bf\u662f glm-5.3 \u4e0a\u5347\u4e86 78.7%\uff0c\u4ece 18:00 \u7684 49.21 \u8bcd\u5143/\u79d2\u5347\u81f3 20:00 \u7684 143.23 \u8bcd\u5143/\u79d2\uff1bglm-5.2 \u6ce2\u52a8\u6027\u6700\u5927\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 48.3%\uff0c\u5728 31.0 \u81f3 125.78 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u56e0\u6b64\u5404\u9879\u6570\u636e\u4ee3\u8868\u5b8c\u6574\u7684\u56db\u5c0f\u65f6\u65f6\u95f4\u7a97\u53e3\u3002","total_duration_ns":2294014146,"translated_at":"2026-10-05T21:03:04.948016+00:00","translation_eval_count":216,"translation_model":"glm-5.3","translation_prompt_eval_count":355,"translation_total_duration_ns":850855183,"translation_wall_seconds":1.041,"valid_point_count":96,"wall_seconds":2.605},{"coverage_pct":100.0,"eval_count":589,"generated_at":"2026-10-05T20:03:06.810349+00:00","generated_label":"Oct 05, 2026 \u00b7 20:03 UTC","id":1157,"model":"glm-5.3","period_end":"2026-10-05T20:03:02+00:00","period_label":"Oct 05, 2026 \u00b7 16:03 UTC to Oct 05, 2026 \u00b7 20:03 UTC","period_start":"2026-10-05T16:03:02+00:00","prompt_eval_count":2246,"summary":"- deepseek-v4.1-flash is the strongest model with 138.64 token/s average throughput, peaking at 227.33 token/s at 19:40; nemotron-3-ultra is the weakest at 22.01 token/s average, never exceeding 39.76 token/s across the window.\n- The most operationally significant movement is glm-5.3's late surge, climbing from 49.21 token/s at 18:00 to 143.23 token/s at 20:00 for a 63.4% trend gain, while glm-5.2 shows the highest relative volatility (cv 56.8%) and minimax-m3 dipped to 6.95 token/s at 19:20.\n- No missing-data limitation applies: all 96 expected points are valid (100.0% coverage), and each of the eight models reports 12 of 12 samples over the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model with 138.64 token/s average throughput, peaking at 227.33 token/s at 19:40; nemotron-3-ultra is the weakest at 22.01 token/s average, never exceeding 39.76 token/s across the window.\n- The most operationally significant movement is glm-5.3's late surge, climbing from 49.21 token/s at 18:00 to 143.23 token/s at 20:00 for a 63.4% trend gain, while glm-5.2 shows the highest relative volatility (cv 56.8%) and minimax-m3 dipped to 6.95 token/s at 19:20.\n- No missing-data limitation applies: all 96 expected points are valid (100.0% coverage), and each of the eight models reports 12 of 12 samples over the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model with 138.64 token/s average throughput, peaking at 227.33 token/s at 19:40; nemotron-3-ultra is the weakest at 22.01 token/s average, never exceeding 39.76 token/s across the window.","The most operationally significant movement is glm-5.3's late surge, climbing from 49.21 token/s at 18:00 to 143.23 token/s at 20:00 for a 63.4% trend gain, while glm-5.2 shows the highest relative volatility (cv 56.8%) and minimax-m3 dipped to 6.95 token/s at 19:20.","No missing-data limitation applies: all 96 expected points are valid (100.0% coverage), and each of the eight models reports 12 of 12 samples over the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 138.64 \u8bcd\u5143/\u79d2\uff0c\u5728 19:40 \u8fbe\u5230\u5cf0\u503c 227.33 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 22.01 \u8bcd\u5143/\u79d2\uff0c\u5728\u6574\u4e2a\u65f6\u95f4\u7a97\u53e3\u5185\u4ece\u672a\u8d85\u8fc7 39.76 \u8bcd\u5143/\u79d2\u3002\n- \u8fd0\u8425\u5c42\u9762\u6700\u663e\u8457\u7684\u53d8\u52a8\u662f glm-5.3 \u7684\u540e\u671f\u6fc0\u589e\uff0c\u4ece 18:00 \u7684 49.21 \u8bcd\u5143/\u79d2\u6500\u5347\u81f3 20:00 \u7684 143.23 \u8bcd\u5143/\u79d2\uff0c\u8d8b\u52bf\u589e\u5e45\u8fbe 63.4%\uff0c\u800c glm-5.2 \u7684\u76f8\u5bf9\u6ce2\u52a8\u6027\u6700\u9ad8\uff08\u53d8\u5f02\u7cfb\u6570 56.8%\uff09\uff0cminimax-m3 \u5728 19:20 \u4e00\u5ea6\u8dcc\u81f3 6.95 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8 96 \u4e2a\u9884\u671f\u6570\u636e\u70b9\u5747\u6709\u6548\uff08\u8986\u76d6\u7387 100.0%\uff09\uff0c\u4e14\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\uff0c\u516b\u4e2a\u6a21\u578b\u4e2d\u7684\u6bcf\u4e00\u4e2a\u90fd\u62a5\u544a\u4e86 12 \u4e2a\u6837\u672c\u4e2d\u7684\u5168\u90e8 12 \u4e2a\u3002","total_duration_ns":2985424030,"translated_at":"2026-10-05T20:03:06.810349+00:00","translation_eval_count":246,"translation_model":"glm-5.3","translation_prompt_eval_count":371,"translation_total_duration_ns":1388717871,"translation_wall_seconds":1.544,"valid_point_count":96,"wall_seconds":3.151},{"coverage_pct":100.0,"eval_count":329,"generated_at":"2026-10-05T19:03:05.371493+00:00","generated_label":"Oct 05, 2026 \u00b7 19:03 UTC","id":1156,"model":"glm-5.3","period_end":"2026-10-05T19:03:01+00:00","period_label":"Oct 05, 2026 \u00b7 15:03 UTC to Oct 05, 2026 \u00b7 19:03 UTC","period_start":"2026-10-05T15:03:01+00:00","prompt_eval_count":2245,"summary":"- deepseek-v4.1-flash is the strongest model at 121.33 token/s average throughput, peaking at 199.80 token/s, while nemotron-3-ultra is the weakest at 23.46 token/s average and never exceeding 39.76 token/s.\n- glm-5.3 shows the most operationally significant volatility, swinging from 146.03 token/s at 15:20 down to 21.87 token/s at 17:00, with a coefficient of variation of 45.0% and a -23.6% trend; deepseek-v4.1-flash also oscillates sharply between 38.85 and 199.80 token/s.\n- No missing-data limitation exists: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 121.33 token/s average throughput, peaking at 199.80 token/s, while nemotron-3-ultra is the weakest at 23.46 token/s average and never exceeding 39.76 token/s.\n- glm-5.3 shows the most operationally significant volatility, swinging from 146.03 token/s at 15:20 down to 21.87 token/s at 17:00, with a coefficient of variation of 45.0% and a -23.6% trend; deepseek-v4.1-flash also oscillates sharply between 38.85 and 199.80 token/s.\n- No missing-data limitation exists: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 121.33 token/s average throughput, peaking at 199.80 token/s, while nemotron-3-ultra is the weakest at 23.46 token/s average and never exceeding 39.76 token/s.","glm-5.3 shows the most operationally significant volatility, swinging from 146.03 token/s at 15:20 down to 21.87 token/s at 17:00, with a coefficient of variation of 45.0% and a -23.6% trend; deepseek-v4.1-flash also oscillates sharply between 38.85 and 199.80 token/s.","No missing-data limitation exists: all eight models report 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 121.33 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 199.80 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 23.46 \u8bcd\u5143/\u79d2\uff0c\u4e14\u4ece\u672a\u8d85\u8fc7 39.76 \u8bcd\u5143/\u79d2\u3002\n- glm-5.3 \u8868\u73b0\u51fa\u6700\u5177\u8fd0\u8425\u610f\u4e49\u7684\u6ce2\u52a8\u6027\uff0c\u4ece 15:20 \u7684 146.03 \u8bcd\u5143/\u79d2\u9aa4\u964d\u81f3 17:00 \u7684 21.87 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 45.0%\uff0c\u8d8b\u52bf\u4e3a -23.6%\uff1bdeepseek-v4.1-flash \u4e5f\u5728 38.85 \u81f3 199.80 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u5267\u70c8\u9707\u8361\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62a5\u544a\u4e86\u9884\u671f\u7684 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u5171\u6709 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":1813271318,"translated_at":"2026-10-05T19:03:05.371493+00:00","translation_eval_count":224,"translation_model":"glm-5.3","translation_prompt_eval_count":356,"translation_total_duration_ns":1257875185,"translation_wall_seconds":1.409,"valid_point_count":96,"wall_seconds":2.153},{"coverage_pct":100.0,"eval_count":683,"generated_at":"2026-10-05T18:03:11.459509+00:00","generated_label":"Oct 05, 2026 \u00b7 18:03 UTC","id":1155,"model":"glm-5.3","period_end":"2026-10-05T18:03:01+00:00","period_label":"Oct 05, 2026 \u00b7 14:03 UTC to Oct 05, 2026 \u00b7 18:03 UTC","period_start":"2026-10-05T14:03:01+00:00","prompt_eval_count":2246,"summary":"- Strongest average throughput was deepseek-v4.1-flash at 111.66 token/s, ahead of gemma4:31b at 103.83 token/s; weakest was nemotron-3-ultra at 24.99 token/s, below minimax-m3 at 43.90 token/s.\n- glm-5.2 showed the most extreme volatility, ranging from 9.97 to 239.87 token/s with a 90.7% coefficient of variation, while glm-5.3 posted the steepest decline, falling 46.4% from 145.53 token/s at 14:20 to 49.21 token/s at 18:00.\n- No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage, though the window covers only four hours.","summary_en":"- Strongest average throughput was deepseek-v4.1-flash at 111.66 token/s, ahead of gemma4:31b at 103.83 token/s; weakest was nemotron-3-ultra at 24.99 token/s, below minimax-m3 at 43.90 token/s.\n- glm-5.2 showed the most extreme volatility, ranging from 9.97 to 239.87 token/s with a 90.7% coefficient of variation, while glm-5.3 posted the steepest decline, falling 46.4% from 145.53 token/s at 14:20 to 49.21 token/s at 18:00.\n- No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage, though the window covers only four hours.","summary_items":["Strongest average throughput was deepseek-v4.1-flash at 111.66 token/s, ahead of gemma4:31b at 103.83 token/s; weakest was nemotron-3-ultra at 24.99 token/s, below minimax-m3 at 43.90 token/s.","glm-5.2 showed the most extreme volatility, ranging from 9.97 to 239.87 token/s with a 90.7% coefficient of variation, while glm-5.3 posted the steepest decline, falling 46.4% from 145.53 token/s at 14:20 to 49.21 token/s at 18:00.","No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage, though the window covers only four hours."],"summary_zh":"- \u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u662f deepseek-v4.1-flash\uff0c\u8fbe 111.66 \u8bcd\u5143/\u79d2\uff0c\u9886\u5148\u4e8e 103.83 \u8bcd\u5143/\u79d2\u7684 gemma4:31b\uff1b\u6700\u5f31\u7684\u662f nemotron-3-ultra\uff0c\u4e3a 24.99 \u8bcd\u5143/\u79d2\uff0c\u4f4e\u4e8e 43.90 \u8bcd\u5143/\u79d2\u7684 minimax-m3\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u6700\u5267\u70c8\u7684\u6ce2\u52a8\u6027\uff0c\u8303\u56f4\u4ece 9.97 \u5230 239.87 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 90.7%\uff0c\u800c glm-5.3 \u51fa\u73b0\u4e86\u6700\u9661\u5ced\u7684\u4e0b\u964d\uff0c\u4ece 14:20 \u7684 145.53 \u8bcd\u5143/\u79d2\u8dcc\u81f3 18:00 \u7684 49.21 \u8bcd\u5143/\u79d2\uff0c\u964d\u5e45\u8fbe 46.4%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u4ea4\u4ed8\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u4ea7\u751f 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\u548c 100.0% \u7684\u8986\u76d6\u7387\uff0c\u4e0d\u8fc7\u8be5\u65f6\u95f4\u7a97\u53e3\u4ec5\u8986\u76d6\u56db\u4e2a\u5c0f\u65f6\u3002","total_duration_ns":7104897550,"translated_at":"2026-10-05T18:03:11.459509+00:00","translation_eval_count":233,"translation_model":"glm-5.3","translation_prompt_eval_count":361,"translation_total_duration_ns":2471508864,"translation_wall_seconds":2.631,"valid_point_count":96,"wall_seconds":7.281},{"coverage_pct":100.0,"eval_count":481,"generated_at":"2026-10-05T17:03:09.262137+00:00","generated_label":"Oct 05, 2026 \u00b7 17:03 UTC","id":1154,"model":"glm-5.3","period_end":"2026-10-05T17:03:01+00:00","period_label":"Oct 05, 2026 \u00b7 13:03 UTC to Oct 05, 2026 \u00b7 17:03 UTC","period_start":"2026-10-05T13:03:01+00:00","prompt_eval_count":2248,"summary":"- glm-5.3 delivered the highest average throughput at 114.93 token/s (peak 154.09 token/s), while nemotron-3-ultra was the weakest at 23.15 token/s average, never exceeding 60.23 token/s.\n- glm-5.2 showed the most operationally significant volatility, swinging between 9.97 and 239.87 token/s with a coefficient of variation of 102.5%; deepseek-v4.1-flash also spiked to 199.8 token/s at 15:40 before falling to 38.85 token/s at 16:40.\n- No missing-data limitation applies: all eight models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- glm-5.3 delivered the highest average throughput at 114.93 token/s (peak 154.09 token/s), while nemotron-3-ultra was the weakest at 23.15 token/s average, never exceeding 60.23 token/s.\n- glm-5.2 showed the most operationally significant volatility, swinging between 9.97 and 239.87 token/s with a coefficient of variation of 102.5%; deepseek-v4.1-flash also spiked to 199.8 token/s at 15:40 before falling to 38.85 token/s at 16:40.\n- No missing-data limitation applies: all eight models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["glm-5.3 delivered the highest average throughput at 114.93 token/s (peak 154.09 token/s), while nemotron-3-ultra was the weakest at 23.15 token/s average, never exceeding 60.23 token/s.","glm-5.2 showed the most operationally significant volatility, swinging between 9.97 and 239.87 token/s with a coefficient of variation of 102.5%; deepseek-v4.1-flash also spiked to 199.8 token/s at 15:40 before falling to 38.85 token/s at 16:40.","No missing-data limitation applies: all eight models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- glm-5.3 \u7684\u5e73\u5747\u541e\u5410\u91cf\u6700\u9ad8\uff0c\u8fbe 114.93 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 154.09 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u8868\u73b0\u6700\u5f31\uff0c\u5e73\u5747\u4ec5 23.15 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 60.23 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u7684\u6ce2\u52a8\u6027\u6700\u5177\u8fd0\u8425\u610f\u4e49\uff0c\u5728 9.97 \u81f3 239.87 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 102.5%\uff1bdeepseek-v4.1-flash \u4e5f\u5728 15:40 \u98d9\u5347\u81f3 199.8 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u5728 16:40 \u56de\u843d\u81f3 38.85 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62a5\u544a 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u5171\u6709 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":2602709810,"translated_at":"2026-10-05T17:03:09.262137+00:00","translation_eval_count":424,"translation_model":"glm-5.3","translation_prompt_eval_count":340,"translation_total_duration_ns":4775288906,"translation_wall_seconds":4.926,"valid_point_count":96,"wall_seconds":2.771},{"coverage_pct":100.0,"eval_count":815,"generated_at":"2026-10-05T16:03:08.246150+00:00","generated_label":"Oct 05, 2026 \u00b7 16:03 UTC","id":1153,"model":"glm-5.3","period_end":"2026-10-05T16:03:01+00:00","period_label":"Oct 05, 2026 \u00b7 12:03 UTC to Oct 05, 2026 \u00b7 16:03 UTC","period_start":"2026-10-05T12:03:01+00:00","prompt_eval_count":2249,"summary":"- Strongest average throughput belongs to deepseek-v4.1-flash at 126.29 token/s (p95 190.37 token/s); weakest is nemotron-3-ultra at 23.42 token/s, peaking at only 60.23 token/s.\n- glm-5.2 shows the most operationally significant volatility, with a 95.3% coefficient of variation, swings between 9.97 and 239.87 token/s, and a -35.1% trend; glm-5.3 rose 48.5% to average 108.11 token/s.\n- No missing-data limitation: all 8 models delivered 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the 12:20\u201316:00 UTC observations.","summary_en":"- Strongest average throughput belongs to deepseek-v4.1-flash at 126.29 token/s (p95 190.37 token/s); weakest is nemotron-3-ultra at 23.42 token/s, peaking at only 60.23 token/s.\n- glm-5.2 shows the most operationally significant volatility, with a 95.3% coefficient of variation, swings between 9.97 and 239.87 token/s, and a -35.1% trend; glm-5.3 rose 48.5% to average 108.11 token/s.\n- No missing-data limitation: all 8 models delivered 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the 12:20\u201316:00 UTC observations.","summary_items":["Strongest average throughput belongs to deepseek-v4.1-flash at 126.29 token/s (p95 190.37 token/s); weakest is nemotron-3-ultra at 23.42 token/s, peaking at only 60.23 token/s.","glm-5.2 shows the most operationally significant volatility, with a 95.3% coefficient of variation, swings between 9.97 and 239.87 token/s, and a -35.1% trend; glm-5.3 rose 48.5% to average 108.11 token/s.","No missing-data limitation: all 8 models delivered 12 of 12 expected samples, 96 valid points, and 100.0% coverage across the 12:20\u201316:00 UTC observations."],"summary_zh":"- \u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u662f deepseek-v4.1-flash\uff0c\u8fbe 126.29 \u8bcd\u5143/\u79d2\uff08p95 \u4e3a 190.37 \u8bcd\u5143/\u79d2\uff09\uff1b\u6700\u5f31\u7684\u662f nemotron-3-ultra\uff0c\u4e3a 23.42 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u4ec5\u4e3a 60.23 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u5bf9\u8fd0\u8425\u5f71\u54cd\u6700\u5927\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 95.3%\uff0c\u5728 9.97 \u81f3 239.87 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\uff0c\u8d8b\u52bf\u4e3a -35.1%\uff1bglm-5.3 \u4e0a\u5347\u4e86 48.5%\uff0c\u5e73\u5747\u8fbe\u5230 108.11 \u8bcd\u5143/\u79d2\u3002\n- \u65e0\u7f3a\u5931\u6570\u636e\u9650\u5236\uff1a\u5168\u90e8 8 \u4e2a\u6a21\u578b\u5747\u4ea4\u4ed8\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5728 12:20\u201316:00 UTC \u89c2\u6d4b\u671f\u95f4\u8986\u76d6\u7387\u8fbe\u5230 100.0%\u3002","total_duration_ns":5089632329,"translated_at":"2026-10-05T16:03:08.246150+00:00","translation_eval_count":207,"translation_model":"glm-5.3","translation_prompt_eval_count":345,"translation_total_duration_ns":1176498676,"translation_wall_seconds":1.509,"valid_point_count":96,"wall_seconds":5.315},{"coverage_pct":100.0,"eval_count":286,"generated_at":"2026-10-05T15:03:05.328181+00:00","generated_label":"Oct 05, 2026 \u00b7 15:03 UTC","id":1152,"model":"glm-5.3","period_end":"2026-10-05T15:03:01+00:00","period_label":"Oct 05, 2026 \u00b7 11:03 UTC to Oct 05, 2026 \u00b7 15:03 UTC","period_start":"2026-10-05T11:03:01+00:00","prompt_eval_count":2250,"summary":"- deepseek-v4.1-flash is the strongest model with an average throughput of 135.61 token/s (maximum 182.65 token/s), while nemotron-3-ultra is the weakest at 21.70 token/s average, peaking at only 60.23 token/s.\n- glm-5.2 shows the most operationally significant volatility, swinging between 9.97 and 239.87 token/s with a coefficient of variation of 85.4%, including a drop to 9.97 token/s at 14:20 followed by a spike to 239.87 token/s at 15:00; deepseek-v4.1-flash also declined 18.2% over the window.\n- No missing-data limitation exists in this dataset: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour period.","summary_en":"- deepseek-v4.1-flash is the strongest model with an average throughput of 135.61 token/s (maximum 182.65 token/s), while nemotron-3-ultra is the weakest at 21.70 token/s average, peaking at only 60.23 token/s.\n- glm-5.2 shows the most operationally significant volatility, swinging between 9.97 and 239.87 token/s with a coefficient of variation of 85.4%, including a drop to 9.97 token/s at 14:20 followed by a spike to 239.87 token/s at 15:00; deepseek-v4.1-flash also declined 18.2% over the window.\n- No missing-data limitation exists in this dataset: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour period.","summary_items":["deepseek-v4.1-flash is the strongest model with an average throughput of 135.61 token/s (maximum 182.65 token/s), while nemotron-3-ultra is the weakest at 21.70 token/s average, peaking at only 60.23 token/s.","glm-5.2 shows the most operationally significant volatility, swinging between 9.97 and 239.87 token/s with a coefficient of variation of 85.4%, including a drop to 9.97 token/s at 14:20 followed by a spike to 239.87 token/s at 15:00; deepseek-v4.1-flash also declined 18.2% over the window.","No missing-data limitation exists in this dataset: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour period."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 135.61 \u8bcd\u5143/\u79d2\uff08\u6700\u9ad8 182.65 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4ec5\u4e3a 21.70 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u4e5f\u4ec5\u6709 60.23 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u5bf9\u8fd0\u8425\u5f71\u54cd\u6700\u5927\u7684\u6ce2\u52a8\u6027\uff0c\u5728 9.97 \u81f3 239.87 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6446\u52a8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 85.4%\uff0c\u5176\u4e2d\u5305\u62ec\u5728 14:20 \u964d\u81f3 9.97 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u5728 15:00 \u98d9\u5347\u81f3 239.87 \u8bcd\u5143/\u79d2\uff1bdeepseek-v4.1-flash \u5728\u8be5\u65f6\u95f4\u7a97\u53e3\u5185\u4e5f\u4e0b\u964d\u4e86 18.2%\u3002\n- \u672c\u6570\u636e\u96c6\u4e0d\u5b58\u5728\u7f3a\u5931\u6570\u636e\u7684\u5c40\u9650\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u8bb0\u5f55\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u4ea7\u751f 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5728\u56db\u5c0f\u65f6\u671f\u95f4\u5185\u8986\u76d6\u7387\u8fbe\u5230 100.0%\u3002","total_duration_ns":1830446583,"translated_at":"2026-10-05T15:03:05.328181+00:00","translation_eval_count":229,"translation_model":"glm-5.3","translation_prompt_eval_count":366,"translation_total_duration_ns":1358028168,"translation_wall_seconds":1.533,"valid_point_count":96,"wall_seconds":2.005},{"coverage_pct":100.0,"eval_count":361,"generated_at":"2026-10-05T14:03:05.630390+00:00","generated_label":"Oct 05, 2026 \u00b7 14:03 UTC","id":1151,"model":"glm-5.3","period_end":"2026-10-05T14:03:01+00:00","period_label":"Oct 05, 2026 \u00b7 10:03 UTC to Oct 05, 2026 \u00b7 14:03 UTC","period_start":"2026-10-05T10:03:01+00:00","prompt_eval_count":2247,"summary":"- deepseek-v4.1-flash is the strongest model with an average throughput of 149.56 token/s (peaking at 212.23 token/s), while nemotron-3-ultra is the weakest at 17.34 token/s average, never exceeding 29.54 token/s.\n- glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 76.8% and swings from 27.27 token/s at 11:00 to 224.82 token/s at 13:40; deepseek-v4-pro also dropped to 11.72 token/s at 14:00.\n- No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model with an average throughput of 149.56 token/s (peaking at 212.23 token/s), while nemotron-3-ultra is the weakest at 17.34 token/s average, never exceeding 29.54 token/s.\n- glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 76.8% and swings from 27.27 token/s at 11:00 to 224.82 token/s at 13:40; deepseek-v4-pro also dropped to 11.72 token/s at 14:00.\n- No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model with an average throughput of 149.56 token/s (peaking at 212.23 token/s), while nemotron-3-ultra is the weakest at 17.34 token/s average, never exceeding 29.54 token/s.","glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 76.8% and swings from 27.27 token/s at 11:00 to 224.82 token/s at 13:40; deepseek-v4-pro also dropped to 11.72 token/s at 14:00.","No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 149.56 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c\u8fbe 212.23 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4ec5 17.34 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 29.54 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u5bf9\u8fd0\u8425\u5f71\u54cd\u6700\u5927\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 76.8%\uff0c\u4ece 11:00 \u7684 27.27 \u8bcd\u5143/\u79d2\u5230 13:40 \u7684 224.82 \u8bcd\u5143/\u79d2\u5927\u5e45\u6ce2\u52a8\uff1bdeepseek-v4-pro \u4e5f\u5728 14:00 \u964d\u81f3 11.72 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u8bb0\u5f55\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u5171\u4ea7\u751f 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":2157480716,"translated_at":"2026-10-05T14:03:05.630390+00:00","translation_eval_count":215,"translation_model":"glm-5.3","translation_prompt_eval_count":346,"translation_total_duration_ns":1300469498,"translation_wall_seconds":1.453,"valid_point_count":96,"wall_seconds":2.496},{"coverage_pct":100.0,"eval_count":271,"generated_at":"2026-10-05T13:03:07.423646+00:00","generated_label":"Oct 05, 2026 \u00b7 13:03 UTC","id":1150,"model":"glm-5.3","period_end":"2026-10-05T13:03:01+00:00","period_label":"Oct 05, 2026 \u00b7 09:03 UTC to Oct 05, 2026 \u00b7 13:03 UTC","period_start":"2026-10-05T09:03:01+00:00","prompt_eval_count":2247,"summary":"- deepseek-v4.1-flash is the strongest model by average throughput at 154.26 token/s, while nemotron-3-ultra is the weakest at 17.55 token/s, roughly nine times lower.\n- glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 81.9% and a trend of +106.1%; it swung from 25.95 token/s at 09:20 to a 193.32 token/s peak at 13:00, including a drop to 45.18 token/s at 12:20 after reaching 177.64 token/s at 12:00.\n- No missing-data limitation exists: all eight models have 12 of 12 expected samples, and the dataset reports 96 valid points with 100.0% coverage over the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model by average throughput at 154.26 token/s, while nemotron-3-ultra is the weakest at 17.55 token/s, roughly nine times lower.\n- glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 81.9% and a trend of +106.1%; it swung from 25.95 token/s at 09:20 to a 193.32 token/s peak at 13:00, including a drop to 45.18 token/s at 12:20 after reaching 177.64 token/s at 12:00.\n- No missing-data limitation exists: all eight models have 12 of 12 expected samples, and the dataset reports 96 valid points with 100.0% coverage over the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model by average throughput at 154.26 token/s, while nemotron-3-ultra is the weakest at 17.55 token/s, roughly nine times lower.","glm-5.2 shows the most operationally significant volatility, with a coefficient of variation of 81.9% and a trend of +106.1%; it swung from 25.95 token/s at 09:20 to a 193.32 token/s peak at 13:00, including a drop to 45.18 token/s at 12:20 after reaching 177.64 token/s at 12:00.","No missing-data limitation exists: all eight models have 12 of 12 expected samples, and the dataset reports 96 valid points with 100.0% coverage over the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u8fbe 154.26 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 17.55 \u8bcd\u5143/\u79d2\uff0c\u4f4e\u4e86\u7ea6\u4e5d\u500d\u3002\n- glm-5.2 \u7684\u6ce2\u52a8\u6027\u6700\u5177\u8fd0\u8425\u610f\u4e49\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 81.9%\uff0c\u8d8b\u52bf\u4e3a +106.1%\uff1b\u5176\u541e\u5410\u91cf\u4ece 09:20 \u7684 25.95 \u8bcd\u5143/\u79d2\u6446\u52a8\u81f3 13:00 \u7684 193.32 \u8bcd\u5143/\u79d2\u5cf0\u503c\uff0c\u5176\u4e2d\u5305\u62ec\u5728 12:00 \u8fbe\u5230 177.64 \u8bcd\u5143/\u79d2\u540e\uff0c\u4e8e 12:20 \u8dcc\u81f3 45.18 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62e5\u6709 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u6570\u636e\u96c6\u62a5\u544a\u4e86 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5728\u56db\u5c0f\u65f6\u65f6\u95f4\u7a97\u53e3\u5185\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":2761068138,"translated_at":"2026-10-05T13:03:07.423646+00:00","translation_eval_count":206,"translation_model":"glm-5.3","translation_prompt_eval_count":353,"translation_total_duration_ns":2290732680,"translation_wall_seconds":2.443,"valid_point_count":96,"wall_seconds":2.989},{"coverage_pct":100.0,"eval_count":674,"generated_at":"2026-10-05T12:03:10.240494+00:00","generated_label":"Oct 05, 2026 \u00b7 12:03 UTC","id":1149,"model":"glm-5.3","period_end":"2026-10-05T12:03:01+00:00","period_label":"Oct 05, 2026 \u00b7 08:03 UTC to Oct 05, 2026 \u00b7 12:03 UTC","period_start":"2026-10-05T08:03:01+00:00","prompt_eval_count":2244,"summary":"- deepseek-v4.1-flash is the strongest model at 152.58 token/s average throughput, while nemotron-3-ultra is the weakest at 17.79 token/s average.\n- The most operationally significant volatility is glm-5.2, whose coefficient of variation is 64.8%; it swung from a 25.95 token/s low to 177.64 token/s at 12:00 UTC, and deepseek-v4-pro dropped sharply to 16.83 token/s in the final sample versus its 93.28 token/s average.\n- No missing-data limitation exists: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 152.58 token/s average throughput, while nemotron-3-ultra is the weakest at 17.79 token/s average.\n- The most operationally significant volatility is glm-5.2, whose coefficient of variation is 64.8%; it swung from a 25.95 token/s low to 177.64 token/s at 12:00 UTC, and deepseek-v4-pro dropped sharply to 16.83 token/s in the final sample versus its 93.28 token/s average.\n- No missing-data limitation exists: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 152.58 token/s average throughput, while nemotron-3-ultra is the weakest at 17.79 token/s average.","The most operationally significant volatility is glm-5.2, whose coefficient of variation is 64.8%; it swung from a 25.95 token/s low to 177.64 token/s at 12:00 UTC, and deepseek-v4-pro dropped sharply to 16.83 token/s in the final sample versus its 93.28 token/s average.","No missing-data limitation exists: all eight models delivered 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6027\u80fd\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 152.58 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u662f\u6027\u80fd\u6700\u5f31\u7684\u6a21\u578b\uff0c\u5e73\u5747\u4e3a 17.79 \u8bcd\u5143/\u79d2\u3002\n- \u8fd0\u884c\u5c42\u9762\u6700\u663e\u8457\u7684\u6ce2\u52a8\u6765\u81ea glm-5.2\uff0c\u5176\u53d8\u5f02\u7cfb\u6570\u4e3a 64.8%\uff1b\u8be5\u6a21\u578b\u4ece 25.95 \u8bcd\u5143/\u79d2\u7684\u4f4e\u70b9\u6ce2\u52a8\u81f3 12:00 UTC \u65f6\u7684 177.64 \u8bcd\u5143/\u79d2\uff0c\u800c deepseek-v4-pro \u5728\u6700\u540e\u4e00\u6b21\u91c7\u6837\u4e2d\u9aa4\u964d\u81f3 16.83 \u8bcd\u5143/\u79d2\uff0c\u8fdc\u4f4e\u4e8e\u5176 93.28 \u8bcd\u5143/\u79d2\u7684\u5e73\u5747\u503c\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u4ea4\u4ed8\u4e86 12 \u4e2a\u9884\u671f\u91c7\u6837\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u5171\u6709 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":6060765700,"translated_at":"2026-10-05T12:03:10.240494+00:00","translation_eval_count":188,"translation_model":"glm-5.3","translation_prompt_eval_count":330,"translation_total_duration_ns":2417711740,"translation_wall_seconds":2.569,"valid_point_count":96,"wall_seconds":6.246},{"coverage_pct":100.0,"eval_count":516,"generated_at":"2026-10-05T11:03:06.052026+00:00","generated_label":"Oct 05, 2026 \u00b7 11:03 UTC","id":1148,"model":"glm-5.3","period_end":"2026-10-05T11:03:01+00:00","period_label":"Oct 05, 2026 \u00b7 07:03 UTC to Oct 05, 2026 \u00b7 11:03 UTC","period_start":"2026-10-05T07:03:01+00:00","prompt_eval_count":2244,"summary":"- deepseek-v4.1-flash is the strongest model at 146.78 token/s average throughput, peaking at 212.23 token/s; nemotron-3-ultra is the weakest at 17.57 token/s average, never exceeding 28.73 token/s.\n- glm-5.3-flash shows the most operationally significant volatility, with a 45.1% coefficient of variation and swings between 53.66 and 232.91 token/s, including a collapse from 232.91 token/s at 08:40 to 53.66 token/s at 09:00; glm-5.3 trended up 42.5%.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.","summary_en":"- deepseek-v4.1-flash is the strongest model at 146.78 token/s average throughput, peaking at 212.23 token/s; nemotron-3-ultra is the weakest at 17.57 token/s average, never exceeding 28.73 token/s.\n- glm-5.3-flash shows the most operationally significant volatility, with a 45.1% coefficient of variation and swings between 53.66 and 232.91 token/s, including a collapse from 232.91 token/s at 08:40 to 53.66 token/s at 09:00; glm-5.3 trended up 42.5%.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.","summary_items":["deepseek-v4.1-flash is the strongest model at 146.78 token/s average throughput, peaking at 212.23 token/s; nemotron-3-ultra is the weakest at 17.57 token/s average, never exceeding 28.73 token/s.","glm-5.3-flash shows the most operationally significant volatility, with a 45.1% coefficient of variation and swings between 53.66 and 232.91 token/s, including a collapse from 232.91 token/s at 08:40 to 53.66 token/s at 09:00; glm-5.3 trended up 42.5%.","No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 146.78 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 212.23 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 17.57 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 28.73 \u8bcd\u5143/\u79d2\u3002\n- glm-5.3-flash \u8868\u73b0\u51fa\u5bf9\u8fd0\u8425\u5f71\u54cd\u6700\u5927\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 45.1%\uff0c\u6ce2\u52a8\u8303\u56f4\u5728 53.66 \u81f3 232.91 \u8bcd\u5143/\u79d2\u4e4b\u95f4\uff0c\u5305\u62ec\u5728 08:40 \u4ece 232.91 \u8bcd\u5143/\u79d2\u9aa4\u964d\u81f3 09:00 \u7684 53.66 \u8bcd\u5143/\u79d2\uff1bglm-5.3 \u5448\u4e0a\u5347\u8d8b\u52bf\uff0c\u6da8\u5e45 42.5%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0cvalid_point_count \u4e3a 96\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u56e0\u6b64\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5f97\u5230\u5b8c\u6574\u5448\u73b0\u3002","total_duration_ns":2674114660,"translated_at":"2026-10-05T11:03:06.052026+00:00","translation_eval_count":217,"translation_model":"glm-5.3","translation_prompt_eval_count":357,"translation_total_duration_ns":1046643841,"translation_wall_seconds":1.198,"valid_point_count":96,"wall_seconds":2.848},{"coverage_pct":100.0,"eval_count":509,"generated_at":"2026-10-05T10:03:06.229918+00:00","generated_label":"Oct 05, 2026 \u00b7 10:03 UTC","id":1147,"model":"glm-5.3","period_end":"2026-10-05T10:03:01+00:00","period_label":"Oct 05, 2026 \u00b7 06:03 UTC to Oct 05, 2026 \u00b7 10:03 UTC","period_start":"2026-10-05T06:03:01+00:00","prompt_eval_count":2244,"summary":"- Strongest average throughput is deepseek-v4.1-flash at 142.14 token/s (p95 176.22 token/s); weakest is nemotron-3-ultra at 18.32 token/s, which never exceeded 28.73 token/s.\n- glm-5.3-flash is the most volatile model, with a 46.0% coefficient of variation and swings between 53.66 and 232.91 token/s; glm-5.3 shows the steepest decline, trending -28.7% from 170.48 token/s at 06:20 to 110.28 token/s at 10:00.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_en":"- Strongest average throughput is deepseek-v4.1-flash at 142.14 token/s (p95 176.22 token/s); weakest is nemotron-3-ultra at 18.32 token/s, which never exceeded 28.73 token/s.\n- glm-5.3-flash is the most volatile model, with a 46.0% coefficient of variation and swings between 53.66 and 232.91 token/s; glm-5.3 shows the steepest decline, trending -28.7% from 170.48 token/s at 06:20 to 110.28 token/s at 10:00.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_items":["Strongest average throughput is deepseek-v4.1-flash at 142.14 token/s (p95 176.22 token/s); weakest is nemotron-3-ultra at 18.32 token/s, which never exceeded 28.73 token/s.","glm-5.3-flash is the most volatile model, with a 46.0% coefficient of variation and swings between 53.66 and 232.91 token/s; glm-5.3 shows the steepest decline, trending -28.7% from 170.48 token/s at 06:20 to 110.28 token/s at 10:00.","No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window."],"summary_zh":"- \u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u662f deepseek-v4.1-flash\uff0c\u8fbe 142.14 \u8bcd\u5143/\u79d2\uff08p95 \u4e3a 176.22 \u8bcd\u5143/\u79d2\uff09\uff1b\u6700\u5f31\u7684\u662f nemotron-3-ultra\uff0c\u4e3a 18.32 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 28.73 \u8bcd\u5143/\u79d2\u3002\n- glm-5.3-flash \u662f\u6ce2\u52a8\u6027\u6700\u5927\u7684\u6a21\u578b\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 46.0%\uff0c\u5728 53.66 \u81f3 232.91 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\uff1bglm-5.3 \u5448\u73b0\u6700\u9661\u5ced\u7684\u4e0b\u964d\u8d8b\u52bf\uff0c\u4ece 06:20 \u7684 170.48 \u8bcd\u5143/\u79d2\u964d\u81f3 10:00 \u7684 110.28 \u8bcd\u5143/\u79d2\uff0c\u8d8b\u52bf\u4e3a -28.7%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5728\u56db\u5c0f\u65f6\u65f6\u95f4\u7a97\u53e3\u5185\u8986\u76d6\u7387\u8fbe\u5230 100.0%\u3002","total_duration_ns":3037303591,"translated_at":"2026-10-05T10:03:06.229918+00:00","translation_eval_count":216,"translation_model":"glm-5.3","translation_prompt_eval_count":347,"translation_total_duration_ns":1059022627,"translation_wall_seconds":1.212,"valid_point_count":96,"wall_seconds":3.212},{"coverage_pct":100.0,"eval_count":301,"generated_at":"2026-10-05T09:03:07.602159+00:00","generated_label":"Oct 05, 2026 \u00b7 09:03 UTC","id":1146,"model":"glm-5.3","period_end":"2026-10-05T09:03:02+00:00","period_label":"Oct 05, 2026 \u00b7 05:03 UTC to Oct 05, 2026 \u00b7 09:03 UTC","period_start":"2026-10-05T05:03:02+00:00","prompt_eval_count":2245,"summary":"- gemma4:31b is the strongest model by average throughput at 129.35 token/s, ahead of deepseek-v4.1-flash at 148.87 token/s... correction: deepseek-v4.1-flash leads at 148.87 token/s, with gemma4:31b second at 129.35 token/s; nemotron-3-ultra is weakest at 21.35 token/s.\n- glm-5.3 shows the sharpest decline, trending down 33.3% from a 170.48 token/s peak at 06:20 to 69.86 token/s at 09:00, while glm-5.3-flash rose 19.7% and spiked to 232.91 token/s at 08:40 before dropping to 53.66 token/s at 09:00.\n- No missing-data limitation exists: all eight models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- gemma4:31b is the strongest model by average throughput at 129.35 token/s, ahead of deepseek-v4.1-flash at 148.87 token/s... correction: deepseek-v4.1-flash leads at 148.87 token/s, with gemma4:31b second at 129.35 token/s; nemotron-3-ultra is weakest at 21.35 token/s.\n- glm-5.3 shows the sharpest decline, trending down 33.3% from a 170.48 token/s peak at 06:20 to 69.86 token/s at 09:00, while glm-5.3-flash rose 19.7% and spiked to 232.91 token/s at 08:40 before dropping to 53.66 token/s at 09:00.\n- No missing-data limitation exists: all eight models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["gemma4:31b is the strongest model by average throughput at 129.35 token/s, ahead of deepseek-v4.1-flash at 148.87 token/s... correction: deepseek-v4.1-flash leads at 148.87 token/s, with gemma4:31b second at 129.35 token/s; nemotron-3-ultra is weakest at 21.35 token/s.","glm-5.3 shows the sharpest decline, trending down 33.3% from a 170.48 token/s peak at 06:20 to 69.86 token/s at 09:00, while glm-5.3-flash rose 19.7% and spiked to 232.91 token/s at 08:40 before dropping to 53.66 token/s at 09:00.","No missing-data limitation exists: all eight models have 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- gemma4:31b \u662f\u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u4e3a 129.35 \u8bcd\u5143/\u79d2\uff0c\u9886\u5148\u4e8e deepseek-v4.1-flash \u7684 148.87 \u8bcd\u5143/\u79d2\u2026\u2026\u66f4\u6b63\uff1adeepseek-v4.1-flash \u4ee5 148.87 \u8bcd\u5143/\u79d2\u9886\u5148\uff0cgemma4:31b \u4ee5 129.35 \u8bcd\u5143/\u79d2\u4f4d\u5c45\u7b2c\u4e8c\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 21.35 \u8bcd\u5143/\u79d2\u3002\n- glm-5.3 \u964d\u5e45\u6700\u4e3a\u5267\u70c8\uff0c\u4ece 06:20 \u7684 170.48 \u8bcd\u5143/\u79d2\u5cf0\u503c\u4e0b\u964d 33.3% \u81f3 09:00 \u7684 69.86 \u8bcd\u5143/\u79d2\uff0c\u800c glm-5.3-flash \u4e0a\u5347\u4e86 19.7%\uff0c\u5e76\u5728 08:40 \u98d9\u5347\u81f3 232.91 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u5728 09:00 \u56de\u843d\u81f3 53.66 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62e5\u6709 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u5171\u6709 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":2782836094,"translated_at":"2026-10-05T09:03:07.602159+00:00","translation_eval_count":260,"translation_model":"glm-5.3","translation_prompt_eval_count":388,"translation_total_duration_ns":2358920167,"translation_wall_seconds":2.547,"valid_point_count":96,"wall_seconds":2.958},{"coverage_pct":100.0,"eval_count":314,"generated_at":"2026-10-05T08:03:05.127450+00:00","generated_label":"Oct 05, 2026 \u00b7 08:03 UTC","id":1145,"model":"glm-5.3","period_end":"2026-10-05T08:03:01+00:00","period_label":"Oct 05, 2026 \u00b7 04:03 UTC to Oct 05, 2026 \u00b7 08:03 UTC","period_start":"2026-10-05T04:03:01+00:00","prompt_eval_count":2244,"summary":"- Strongest average throughput was deepseek-v4.1-flash at 128.78 token/s; weakest was nemotron-3-ultra at 24.57 token/s, roughly one-fifth of the leader's rate.\n- deepseek-v4.1-flash also showed the widest swings, ranging 14.9 to 186.13 token/s with 44.2% coefficient of variation, while glm-5.3 climbed 41.0% overall, peaking at 170.48 token/s before falling back to 112.79 token/s.\n- No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- Strongest average throughput was deepseek-v4.1-flash at 128.78 token/s; weakest was nemotron-3-ultra at 24.57 token/s, roughly one-fifth of the leader's rate.\n- deepseek-v4.1-flash also showed the widest swings, ranging 14.9 to 186.13 token/s with 44.2% coefficient of variation, while glm-5.3 climbed 41.0% overall, peaking at 170.48 token/s before falling back to 112.79 token/s.\n- No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["Strongest average throughput was deepseek-v4.1-flash at 128.78 token/s; weakest was nemotron-3-ultra at 24.57 token/s, roughly one-fifth of the leader's rate.","deepseek-v4.1-flash also showed the widest swings, ranging 14.9 to 186.13 token/s with 44.2% coefficient of variation, while glm-5.3 climbed 41.0% overall, peaking at 170.48 token/s before falling back to 112.79 token/s.","No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, giving 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- \u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u662f deepseek-v4.1-flash\uff0c\u8fbe 128.78 \u8bcd\u5143/\u79d2\uff1b\u6700\u5f31\u7684\u662f nemotron-3-ultra\uff0c\u4e3a 24.57 \u8bcd\u5143/\u79d2\uff0c\u7ea6\u4e3a\u9886\u5148\u8005\u901f\u7387\u7684\u4e94\u5206\u4e4b\u4e00\u3002\n- deepseek-v4.1-flash \u7684\u6ce2\u52a8\u5e45\u5ea6\u4e5f\u6700\u5927\uff0c\u8303\u56f4\u4ece 14.9 \u5230 186.13 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 44.2%\uff1b\u800c glm-5.3 \u6574\u4f53\u6500\u5347 41.0%\uff0c\u5728\u56de\u843d\u81f3 112.79 \u8bcd\u5143/\u79d2\u4e4b\u524d\u8fbe\u5230 170.48 \u8bcd\u5143/\u79d2\u7684\u5cf0\u503c\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u4ea4\u4ed8\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u5171\u4ea7\u751f 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":2143078855,"translated_at":"2026-10-05T08:03:05.127450+00:00","translation_eval_count":203,"translation_model":"glm-5.3","translation_prompt_eval_count":328,"translation_total_duration_ns":1191823798,"translation_wall_seconds":1.347,"valid_point_count":96,"wall_seconds":2.307},{"coverage_pct":100.0,"eval_count":514,"generated_at":"2026-10-05T07:03:08.988982+00:00","generated_label":"Oct 05, 2026 \u00b7 07:03 UTC","id":1144,"model":"glm-5.3","period_end":"2026-10-05T07:03:01+00:00","period_label":"Oct 05, 2026 \u00b7 03:03 UTC to Oct 05, 2026 \u00b7 07:03 UTC","period_start":"2026-10-05T03:03:01+00:00","prompt_eval_count":2244,"summary":"- deepseek-v4.1-flash is the strongest model at 135.35 token/s average throughput, peaking at 209.48 token/s; nemotron-3-ultra is the weakest at 23.57 token/s average, never exceeding 46.49 token/s.\n- The most operationally significant movement is deepseek-v4.1-flash's 52.5% upward trend, including a dip to 14.9 token/s at 05:00 before recovering to 165.38 token/s by 07:00; glm-5.2 shows the highest volatility at 65.6% coefficient of variation.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.","summary_en":"- deepseek-v4.1-flash is the strongest model at 135.35 token/s average throughput, peaking at 209.48 token/s; nemotron-3-ultra is the weakest at 23.57 token/s average, never exceeding 46.49 token/s.\n- The most operationally significant movement is deepseek-v4.1-flash's 52.5% upward trend, including a dip to 14.9 token/s at 05:00 before recovering to 165.38 token/s by 07:00; glm-5.2 shows the highest volatility at 65.6% coefficient of variation.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.","summary_items":["deepseek-v4.1-flash is the strongest model at 135.35 token/s average throughput, peaking at 209.48 token/s; nemotron-3-ultra is the weakest at 23.57 token/s average, never exceeding 46.49 token/s.","The most operationally significant movement is deepseek-v4.1-flash's 52.5% upward trend, including a dip to 14.9 token/s at 05:00 before recovering to 165.38 token/s by 07:00; glm-5.2 shows the highest volatility at 65.6% coefficient of variation.","No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 135.35 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 209.48 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 23.57 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 46.49 \u8bcd\u5143/\u79d2\u3002\n- \u8fd0\u8425\u5c42\u9762\u6700\u663e\u8457\u7684\u53d8\u52a8\u662f deepseek-v4.1-flash \u7684 52.5% \u4e0a\u5347\u8d8b\u52bf\uff0c\u5176\u4e2d\u5305\u62ec\u5728 05:00 \u964d\u81f3 14.9 \u8bcd\u5143/\u79d2\u7684\u4f4e\u8c37\uff0c\u968f\u540e\u5728 07:00 \u524d\u6062\u590d\u81f3 165.38 \u8bcd\u5143/\u79d2\uff1bglm-5.2 \u7684\u6ce2\u52a8\u6027\u6700\u9ad8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 65.6%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12/12 \u4e2a\u6837\u672c\uff0cvalid_point_count \u4e3a 96\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u56e0\u6b64\u8be5\u56db\u5c0f\u65f6\u65f6\u95f4\u7a97\u53e3\u5f97\u5230\u5b8c\u6574\u5448\u73b0\u3002","total_duration_ns":4724176506,"translated_at":"2026-10-05T07:03:08.988982+00:00","translation_eval_count":196,"translation_model":"glm-5.3","translation_prompt_eval_count":348,"translation_total_duration_ns":2392033694,"translation_wall_seconds":2.538,"valid_point_count":96,"wall_seconds":4.89},{"coverage_pct":100.0,"eval_count":323,"generated_at":"2026-10-05T05:03:05.356335+00:00","generated_label":"Oct 05, 2026 \u00b7 05:03 UTC","id":1143,"model":"glm-5.3","period_end":"2026-10-05T05:03:01+00:00","period_label":"Oct 05, 2026 \u00b7 01:03 UTC to Oct 05, 2026 \u00b7 05:03 UTC","period_start":"2026-10-05T01:03:01+00:00","prompt_eval_count":2246,"summary":"- Highest average throughput was deepseek-v4.1-flash at 122.44 token/s, edging glm-5.3-flash (122.37 token/s) and gemma4:31b (119.24 token/s); lowest was nemotron-3-ultra at 30.60 token/s, below minimax-m3's 37.69 token/s.\n- glm-5.2 showed the sharpest volatility, with a coefficient of variation of 82.8% and swings between a 202.79 token/s spike at 02:00 and a 21.44 token/s low at 04:40; deepseek-v4-pro trended up 53.1% while glm-5.3-flash fell 24.6%.\n- No missing-data limitation applies: all eight models delivered 12 of 12 samples, with 96 valid points and 100.0% coverage, though the window spans only four hours.","summary_en":"- Highest average throughput was deepseek-v4.1-flash at 122.44 token/s, edging glm-5.3-flash (122.37 token/s) and gemma4:31b (119.24 token/s); lowest was nemotron-3-ultra at 30.60 token/s, below minimax-m3's 37.69 token/s.\n- glm-5.2 showed the sharpest volatility, with a coefficient of variation of 82.8% and swings between a 202.79 token/s spike at 02:00 and a 21.44 token/s low at 04:40; deepseek-v4-pro trended up 53.1% while glm-5.3-flash fell 24.6%.\n- No missing-data limitation applies: all eight models delivered 12 of 12 samples, with 96 valid points and 100.0% coverage, though the window spans only four hours.","summary_items":["Highest average throughput was deepseek-v4.1-flash at 122.44 token/s, edging glm-5.3-flash (122.37 token/s) and gemma4:31b (119.24 token/s); lowest was nemotron-3-ultra at 30.60 token/s, below minimax-m3's 37.69 token/s.","glm-5.2 showed the sharpest volatility, with a coefficient of variation of 82.8% and swings between a 202.79 token/s spike at 02:00 and a 21.44 token/s low at 04:40; deepseek-v4-pro trended up 53.1% while glm-5.3-flash fell 24.6%.","No missing-data limitation applies: all eight models delivered 12 of 12 samples, with 96 valid points and 100.0% coverage, though the window spans only four hours."],"summary_zh":"- \u5e73\u5747\u541e\u5410\u91cf\u6700\u9ad8\u7684\u662f deepseek-v4.1-flash\uff0c\u8fbe 122.44 \u8bcd\u5143/\u79d2\uff0c\u7565\u9ad8\u4e8e glm-5.3-flash\uff08122.37 \u8bcd\u5143/\u79d2\uff09\u548c gemma4:31b\uff08119.24 \u8bcd\u5143/\u79d2\uff09\uff1b\u6700\u4f4e\u7684\u662f nemotron-3-ultra\uff0c\u4e3a 30.60 \u8bcd\u5143/\u79d2\uff0c\u4f4e\u4e8e minimax-m3 \u7684 37.69 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u7684\u6ce2\u52a8\u6027\u6700\u4e3a\u5267\u70c8\uff0c\u53d8\u5f02\u7cfb\u6570\u8fbe 82.8%\uff0c\u5728 02:00 \u51fa\u73b0 202.79 \u8bcd\u5143/\u79d2\u7684\u5cf0\u503c\u4e0e 04:40 \u51fa\u73b0 21.44 \u8bcd\u5143/\u79d2\u7684\u4f4e\u8c37\u4e4b\u95f4\u5927\u5e45\u6ce2\u52a8\uff1bdeepseek-v4-pro \u5448\u4e0a\u5347\u8d8b\u52bf\uff0c\u6da8\u5e45 53.1%\uff0c\u800c glm-5.3-flash \u4e0b\u964d 24.6%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u63d0\u4f9b\u4e86 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5171\u6709 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u4e0d\u8fc7\u8be5\u65f6\u95f4\u7a97\u53e3\u4ec5\u8de8\u8d8a\u56db\u4e2a\u5c0f\u65f6\u3002","total_duration_ns":2025304575,"translated_at":"2026-10-05T05:03:05.356335+00:00","translation_eval_count":245,"translation_model":"glm-5.3","translation_prompt_eval_count":372,"translation_total_duration_ns":917575979,"translation_wall_seconds":1.074,"valid_point_count":96,"wall_seconds":2.37},{"coverage_pct":100.0,"eval_count":444,"generated_at":"2026-10-05T04:03:05.500396+00:00","generated_label":"Oct 05, 2026 \u00b7 04:03 UTC","id":1142,"model":"glm-5.3","period_end":"2026-10-05T04:03:02+00:00","period_label":"Oct 05, 2026 \u00b7 00:03 UTC to Oct 05, 2026 \u00b7 04:03 UTC","period_start":"2026-10-05T00:03:02+00:00","prompt_eval_count":2248,"summary":"- Highest average throughput was gemma4:31b at 133.1 token/s across 12 samples (range 81.25 to 170.9 token/s), while nemotron-3-ultra was lowest at 27.93 token/s, never exceeding 54.07 token/s.\n- glm-5.2 showed the most operationally significant volatility, with a 70.5% coefficient of variation, swings between 24.83 and 202.79 token/s, and a -31.2% trend; deepseek-v4.1-flash also swung between 24.14 and 226.05 token/s.\n- No missing-data limitation applies: all 8 models reported 12 of 12 expected samples, giving 96 valid points and 100.0% coverage over the four-hour window.","summary_en":"- Highest average throughput was gemma4:31b at 133.1 token/s across 12 samples (range 81.25 to 170.9 token/s), while nemotron-3-ultra was lowest at 27.93 token/s, never exceeding 54.07 token/s.\n- glm-5.2 showed the most operationally significant volatility, with a 70.5% coefficient of variation, swings between 24.83 and 202.79 token/s, and a -31.2% trend; deepseek-v4.1-flash also swung between 24.14 and 226.05 token/s.\n- No missing-data limitation applies: all 8 models reported 12 of 12 expected samples, giving 96 valid points and 100.0% coverage over the four-hour window.","summary_items":["Highest average throughput was gemma4:31b at 133.1 token/s across 12 samples (range 81.25 to 170.9 token/s), while nemotron-3-ultra was lowest at 27.93 token/s, never exceeding 54.07 token/s.","glm-5.2 showed the most operationally significant volatility, with a 70.5% coefficient of variation, swings between 24.83 and 202.79 token/s, and a -31.2% trend; deepseek-v4.1-flash also swung between 24.14 and 226.05 token/s.","No missing-data limitation applies: all 8 models reported 12 of 12 expected samples, giving 96 valid points and 100.0% coverage over the four-hour window."],"summary_zh":"- \u5e73\u5747\u541e\u5410\u91cf\u6700\u9ad8\u7684\u662f gemma4:31b\uff0c\u5728 12 \u4e2a\u6837\u672c\u4e2d\u8fbe\u5230 133.1 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4\u4ece 81.25 \u5230 170.9 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u4f4e\uff0c\u4e3a 27.93 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 54.07 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u8868\u73b0\u51fa\u5bf9\u8fd0\u884c\u5f71\u54cd\u6700\u663e\u8457\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 70.5%\uff0c\u5728 24.83 \u548c 202.79 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6446\u52a8\uff0c\u8d8b\u52bf\u4e3a -31.2%\uff1bdeepseek-v4.1-flash \u4e5f\u5728 24.14 \u548c 226.05 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6446\u52a8\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8 8 \u4e2a\u6a21\u578b\u5747\u62a5\u544a\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u63d0\u4f9b 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":1952844056,"translated_at":"2026-10-05T04:03:05.500396+00:00","translation_eval_count":215,"translation_model":"glm-5.3","translation_prompt_eval_count":345,"translation_total_duration_ns":1041588283,"translation_wall_seconds":1.363,"valid_point_count":96,"wall_seconds":2.121},{"coverage_pct":100.0,"eval_count":964,"generated_at":"2026-10-05T03:03:11.927006+00:00","generated_label":"Oct 05, 2026 \u00b7 03:03 UTC","id":1141,"model":"glm-5.3","period_end":"2026-10-05T03:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 23:03 UTC to Oct 05, 2026 \u00b7 03:03 UTC","period_start":"2026-10-04T23:03:01+00:00","prompt_eval_count":2248,"summary":"- deepseek-v4.1-flash is the strongest model at 150.73 token/s average throughput, ahead of gemma4:31b at 132.85 token/s and glm-5.3-flash at 131.83 token/s; nemotron-3-ultra is the weakest at 39.67 token/s average, below minimax-m3 at 48.47 token/s.\n- glm-5.2 shows the most volatility, with a 74.2% coefficient of variation, ranging from 24.52 to 202.79 token/s, including a spike to 202.79 token/s at 02:00 UTC; deepseek-v4.1-flash also swings sharply, dropping to 24.14 token/s at 01:20 after peaking at 226.05 token/s at 00:40.\n- No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 150.73 token/s average throughput, ahead of gemma4:31b at 132.85 token/s and glm-5.3-flash at 131.83 token/s; nemotron-3-ultra is the weakest at 39.67 token/s average, below minimax-m3 at 48.47 token/s.\n- glm-5.2 shows the most volatility, with a 74.2% coefficient of variation, ranging from 24.52 to 202.79 token/s, including a spike to 202.79 token/s at 02:00 UTC; deepseek-v4.1-flash also swings sharply, dropping to 24.14 token/s at 01:20 after peaking at 226.05 token/s at 00:40.\n- No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 150.73 token/s average throughput, ahead of gemma4:31b at 132.85 token/s and glm-5.3-flash at 131.83 token/s; nemotron-3-ultra is the weakest at 39.67 token/s average, below minimax-m3 at 48.47 token/s.","glm-5.2 shows the most volatility, with a 74.2% coefficient of variation, ranging from 24.52 to 202.79 token/s, including a spike to 202.79 token/s at 02:00 UTC; deepseek-v4.1-flash also swings sharply, dropping to 24.14 token/s at 01:20 after peaking at 226.05 token/s at 00:40.","No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 150.73 \u8bcd\u5143/\u79d2\uff0c\u9886\u5148\u4e8e gemma4:31b \u7684 132.85 \u8bcd\u5143/\u79d2\u548c glm-5.3-flash \u7684 131.83 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 39.67 \u8bcd\u5143/\u79d2\uff0c\u4f4e\u4e8e minimax-m3 \u7684 48.47 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u7684\u6ce2\u52a8\u6027\u6700\u5927\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 74.2%\uff0c\u8303\u56f4\u4ece 24.52 \u5230 202.79 \u8bcd\u5143/\u79d2\uff0c\u5176\u4e2d\u5305\u62ec\u5728 02:00 UTC \u98d9\u5347\u81f3 202.79 \u8bcd\u5143/\u79d2\uff1bdeepseek-v4.1-flash \u540c\u6837\u6ce2\u52a8\u5267\u70c8\uff0c\u5728 00:40 \u8fbe\u5230\u5cf0\u503c 226.05 \u8bcd\u5143/\u79d2\u540e\uff0c\u4e8e 01:20 \u8dcc\u81f3 24.14 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0cvalid_point_count \u4e3a 96\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":7169793104,"translated_at":"2026-10-05T03:03:11.927006+00:00","translation_eval_count":249,"translation_model":"glm-5.3","translation_prompt_eval_count":390,"translation_total_duration_ns":2509414243,"translation_wall_seconds":2.651,"valid_point_count":96,"wall_seconds":7.33},{"coverage_pct":100.0,"eval_count":338,"generated_at":"2026-10-05T02:03:08.194849+00:00","generated_label":"Oct 05, 2026 \u00b7 02:03 UTC","id":1140,"model":"glm-5.3","period_end":"2026-10-05T02:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 22:03 UTC to Oct 05, 2026 \u00b7 02:03 UTC","period_start":"2026-10-04T22:03:01+00:00","prompt_eval_count":2249,"summary":"- deepseek-v4.1-flash is the strongest model by average throughput at 140.56 token/s, narrowly ahead of gemma4:31b at 140.37 token/s; nemotron-3-ultra is the weakest at 44.39 token/s, below minimax-m3 at 55.41 token/s.\n- glm-5.2 shows the sharpest swing, rising 123.0% over the window to a peak of 202.79 token/s at 02:00 after dipping to 24.52 token/s at 00:00, with the highest volatility at 72.8% coefficient of variation; minimax-m3 fell 48.6% to 11.73 token/s.\n- No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, and the dataset records 96 valid points at 100.0% coverage.","summary_en":"- deepseek-v4.1-flash is the strongest model by average throughput at 140.56 token/s, narrowly ahead of gemma4:31b at 140.37 token/s; nemotron-3-ultra is the weakest at 44.39 token/s, below minimax-m3 at 55.41 token/s.\n- glm-5.2 shows the sharpest swing, rising 123.0% over the window to a peak of 202.79 token/s at 02:00 after dipping to 24.52 token/s at 00:00, with the highest volatility at 72.8% coefficient of variation; minimax-m3 fell 48.6% to 11.73 token/s.\n- No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, and the dataset records 96 valid points at 100.0% coverage.","summary_items":["deepseek-v4.1-flash is the strongest model by average throughput at 140.56 token/s, narrowly ahead of gemma4:31b at 140.37 token/s; nemotron-3-ultra is the weakest at 44.39 token/s, below minimax-m3 at 55.41 token/s.","glm-5.2 shows the sharpest swing, rising 123.0% over the window to a peak of 202.79 token/s at 02:00 after dipping to 24.52 token/s at 00:00, with the highest volatility at 72.8% coefficient of variation; minimax-m3 fell 48.6% to 11.73 token/s.","No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, and the dataset records 96 valid points at 100.0% coverage."],"summary_zh":"- deepseek-v4.1-flash \u662f\u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u8fbe\u5230 140.56 \u8bcd\u5143/\u79d2\uff0c\u7565\u5fae\u9886\u5148\u4e8e gemma4:31b \u7684 140.37 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u4e3a 44.39 \u8bcd\u5143/\u79d2\uff0c\u4f4e\u4e8e minimax-m3 \u7684 55.41 \u8bcd\u5143/\u79d2\u3002\n- glm-5.2 \u6ce2\u52a8\u6700\u4e3a\u5267\u70c8\uff0c\u5728\u7a97\u53e3\u671f\u5185\u4e0a\u5347 123.0%\uff0c\u5728 00:00 \u8dcc\u81f3 24.52 \u8bcd\u5143/\u79d2\u540e\uff0c\u4e8e 02:00 \u8fbe\u5230 202.79 \u8bcd\u5143/\u79d2\u7684\u5cf0\u503c\uff0c\u53d8\u5f02\u7cfb\u6570\u9ad8\u8fbe 72.8%\uff0c\u6ce2\u52a8\u6027\u6700\u9ad8\uff1bminimax-m3 \u4e0b\u8dcc 48.6%\uff0c\u964d\u81f3 11.73 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u4ea4\u4ed8\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u6570\u636e\u96c6\u8bb0\u5f55\u4e86 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":3982320593,"translated_at":"2026-10-05T02:03:08.194849+00:00","translation_eval_count":226,"translation_model":"glm-5.3","translation_prompt_eval_count":362,"translation_total_duration_ns":2316450491,"translation_wall_seconds":2.466,"valid_point_count":96,"wall_seconds":4.148},{"coverage_pct":100.0,"eval_count":579,"generated_at":"2026-10-05T01:03:09.212148+00:00","generated_label":"Oct 05, 2026 \u00b7 01:03 UTC","id":1139,"model":"glm-5.3","period_end":"2026-10-05T01:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 21:03 UTC to Oct 05, 2026 \u00b7 01:03 UTC","period_start":"2026-10-04T21:03:01+00:00","prompt_eval_count":2246,"summary":"- gemma4:31b is the strongest model by average throughput at 143.44 token/s, ahead of deepseek-v4.1-flash at 135.52 token/s; nemotron-3-ultra is the weakest at 46.7 token/s average, below minimax-m3 at 62.55 token/s.\n- The most operationally significant volatility is glm-5.3, which swings between 19.0 and 180.67 token/s (cv 50.8%), including a drop to 19.0 token/s at 00:00; deepseek-v4.1-flash shows the steepest upward trend at +52.8%, peaking at 226.05 token/s at 00:40.\n- No missing-data limitation applies: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.","summary_en":"- gemma4:31b is the strongest model by average throughput at 143.44 token/s, ahead of deepseek-v4.1-flash at 135.52 token/s; nemotron-3-ultra is the weakest at 46.7 token/s average, below minimax-m3 at 62.55 token/s.\n- The most operationally significant volatility is glm-5.3, which swings between 19.0 and 180.67 token/s (cv 50.8%), including a drop to 19.0 token/s at 00:00; deepseek-v4.1-flash shows the steepest upward trend at +52.8%, peaking at 226.05 token/s at 00:40.\n- No missing-data limitation applies: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.","summary_items":["gemma4:31b is the strongest model by average throughput at 143.44 token/s, ahead of deepseek-v4.1-flash at 135.52 token/s; nemotron-3-ultra is the weakest at 46.7 token/s average, below minimax-m3 at 62.55 token/s.","The most operationally significant volatility is glm-5.3, which swings between 19.0 and 180.67 token/s (cv 50.8%), including a drop to 19.0 token/s at 00:00; deepseek-v4.1-flash shows the steepest upward trend at +52.8%, peaking at 226.05 token/s at 00:40.","No missing-data limitation applies: all 8 models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented."],"summary_zh":"- gemma4:31b \u662f\u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u8fbe\u5230 143.44 \u8bcd\u5143/\u79d2\uff0c\u9886\u5148\u4e8e deepseek-v4.1-flash \u7684 135.52 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 46.7 \u8bcd\u5143/\u79d2\uff0c\u4f4e\u4e8e minimax-m3 \u7684 62.55 \u8bcd\u5143/\u79d2\u3002\n- \u8fd0\u8425\u4e0a\u6700\u663e\u8457\u7684\u6ce2\u52a8\u6027\u6765\u81ea glm-5.3\uff0c\u5176\u5728 19.0 \u81f3 180.67 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\uff08cv 50.8%\uff09\uff0c\u5305\u62ec\u5728 00:00 \u964d\u81f3 19.0 \u8bcd\u5143/\u79d2\uff1bdeepseek-v4.1-flash \u5448\u73b0\u6700\u9661\u5ced\u7684\u4e0a\u5347\u8d8b\u52bf\uff0c\u8fbe +52.8%\uff0c\u5728 00:40 \u8fbe\u5230\u5cf0\u503c 226.05 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8 8 \u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0cvalid_point_count \u4e3a 96\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u56e0\u6b64\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5f97\u5230\u5b8c\u6574\u5448\u73b0\u3002","total_duration_ns":5907871081,"translated_at":"2026-10-05T01:03:09.212148+00:00","translation_eval_count":234,"translation_model":"glm-5.3","translation_prompt_eval_count":374,"translation_total_duration_ns":1015900882,"translation_wall_seconds":1.16,"valid_point_count":96,"wall_seconds":6.072},{"coverage_pct":100.0,"eval_count":486,"generated_at":"2026-10-05T00:03:05.215781+00:00","generated_label":"Oct 05, 2026 \u00b7 00:03 UTC","id":1138,"model":"glm-5.3","period_end":"2026-10-05T00:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 20:03 UTC to Oct 05, 2026 \u00b7 00:03 UTC","period_start":"2026-10-04T20:03:01+00:00","prompt_eval_count":2244,"summary":"- gemma4:31b delivered the strongest average throughput at 129.49 token/s with a peak of 173.62 token/s, while nemotron-3-ultra was weakest at 47.04 token/s average, never exceeding 94.23 token/s.\n- glm-5.3 showed the most operationally significant volatility, swinging between 19.0 and 180.67 token/s with a 45.6% coefficient of variation; deepseek-v4-pro also fell sharply from 112.6 token/s at 23:20 to 12.59 token/s at 23:40.\n- No missing-data limitation applies: all eight models recorded 12 of 12 samples, giving 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- gemma4:31b delivered the strongest average throughput at 129.49 token/s with a peak of 173.62 token/s, while nemotron-3-ultra was weakest at 47.04 token/s average, never exceeding 94.23 token/s.\n- glm-5.3 showed the most operationally significant volatility, swinging between 19.0 and 180.67 token/s with a 45.6% coefficient of variation; deepseek-v4-pro also fell sharply from 112.6 token/s at 23:20 to 12.59 token/s at 23:40.\n- No missing-data limitation applies: all eight models recorded 12 of 12 samples, giving 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["gemma4:31b delivered the strongest average throughput at 129.49 token/s with a peak of 173.62 token/s, while nemotron-3-ultra was weakest at 47.04 token/s average, never exceeding 94.23 token/s.","glm-5.3 showed the most operationally significant volatility, swinging between 19.0 and 180.67 token/s with a 45.6% coefficient of variation; deepseek-v4-pro also fell sharply from 112.6 token/s at 23:20 to 12.59 token/s at 23:40.","No missing-data limitation applies: all eight models recorded 12 of 12 samples, giving 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- gemma4:31b \u4ee5 129.49 \u8bcd\u5143/\u79d2\u7684\u5e73\u5747\u541e\u5410\u91cf\u8868\u73b0\u6700\u5f3a\uff0c\u5cf0\u503c\u8fbe 173.62 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4ec5\u4e3a 47.04 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 94.23 \u8bcd\u5143/\u79d2\u3002\n- glm-5.3 \u8868\u73b0\u51fa\u5bf9\u8fd0\u8425\u5f71\u54cd\u6700\u5927\u7684\u6ce2\u52a8\u6027\uff0c\u5728 19.0 \u81f3 180.67 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6446\u52a8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 45.6%\uff1bdeepseek-v4-pro \u4e5f\u4ece 23:20 \u7684 112.6 \u8bcd\u5143/\u79d2\u9aa4\u964d\u81f3 23:40 \u7684 12.59 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u8bb0\u5f55\u4e86 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u5171\u4ea7\u751f 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":2323650883,"translated_at":"2026-10-05T00:03:05.215781+00:00","translation_eval_count":226,"translation_model":"glm-5.3","translation_prompt_eval_count":338,"translation_total_duration_ns":974150077,"translation_wall_seconds":1.123,"valid_point_count":96,"wall_seconds":2.487},{"coverage_pct":100.0,"eval_count":541,"generated_at":"2026-10-04T23:03:08.864761+00:00","generated_label":"Oct 04, 2026 \u00b7 23:03 UTC","id":1137,"model":"glm-5.3","period_end":"2026-10-04T23:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 19:03 UTC to Oct 04, 2026 \u00b7 23:03 UTC","period_start":"2026-10-04T19:03:01+00:00","prompt_eval_count":2244,"summary":"- glm-5.3 is the strongest model at 129.58 token/s average output throughput, ahead of gemma4:31b at 124.19 token/s; nemotron-3-ultra is the weakest at 37.83 token/s average, never exceeding 75.69 token/s.\n- deepseek-v4.1-flash shows the most operationally significant volatility, with a 60.6% coefficient of variation and swings from 19.01 to 185.99 token/s; glm-5.3 also declined 22.0% across the window, ending at 63.23 token/s versus its 178.26 token/s peak.\n- No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage, though the four-hour window restricts longer-term assessment.","summary_en":"- glm-5.3 is the strongest model at 129.58 token/s average output throughput, ahead of gemma4:31b at 124.19 token/s; nemotron-3-ultra is the weakest at 37.83 token/s average, never exceeding 75.69 token/s.\n- deepseek-v4.1-flash shows the most operationally significant volatility, with a 60.6% coefficient of variation and swings from 19.01 to 185.99 token/s; glm-5.3 also declined 22.0% across the window, ending at 63.23 token/s versus its 178.26 token/s peak.\n- No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage, though the four-hour window restricts longer-term assessment.","summary_items":["glm-5.3 is the strongest model at 129.58 token/s average output throughput, ahead of gemma4:31b at 124.19 token/s; nemotron-3-ultra is the weakest at 37.83 token/s average, never exceeding 75.69 token/s.","deepseek-v4.1-flash shows the most operationally significant volatility, with a 60.6% coefficient of variation and swings from 19.01 to 185.99 token/s; glm-5.3 also declined 22.0% across the window, ending at 63.23 token/s versus its 178.26 token/s peak.","No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, totaling 96 valid points at 100.0% coverage, though the four-hour window restricts longer-term assessment."],"summary_zh":"- glm-5.3 \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u8f93\u51fa\u541e\u5410\u91cf\u4e3a 129.58 \u8bcd\u5143/\u79d2\uff0c\u9886\u5148\u4e8e 124.19 \u8bcd\u5143/\u79d2\u7684 gemma4:31b\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 37.83 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 75.69 \u8bcd\u5143/\u79d2\u3002\n- deepseek-v4.1-flash \u8868\u73b0\u51fa\u5bf9\u8fd0\u8425\u5f71\u54cd\u6700\u5927\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 60.6%\uff0c\u6ce2\u52a8\u8303\u56f4\u4ece 19.01 \u5230 185.99 \u8bcd\u5143/\u79d2\uff1bglm-5.3 \u5728\u6574\u4e2a\u65f6\u95f4\u7a97\u53e3\u5185\u4e5f\u4e0b\u964d\u4e86 22.0%\uff0c\u6700\u7ec8\u4e3a 63.23 \u8bcd\u5143/\u79d2\uff0c\u800c\u5176\u5cf0\u503c\u4e3a 178.26 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8 8 \u4e2a\u6a21\u578b\u5747\u4ea4\u4ed8\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5171\u8ba1 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u4f46\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u9650\u5236\u4e86\u5bf9\u66f4\u957f\u671f\u60c5\u51b5\u7684\u8bc4\u4f30\u3002","total_duration_ns":4834925072,"translated_at":"2026-10-04T23:03:08.864761+00:00","translation_eval_count":213,"translation_model":"glm-5.3","translation_prompt_eval_count":357,"translation_total_duration_ns":2575173107,"translation_wall_seconds":2.718,"valid_point_count":96,"wall_seconds":5.006},{"coverage_pct":100.0,"eval_count":499,"generated_at":"2026-10-04T22:03:07.453144+00:00","generated_label":"Oct 04, 2026 \u00b7 22:03 UTC","id":1136,"model":"glm-5.3","period_end":"2026-10-04T22:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 18:03 UTC to Oct 04, 2026 \u00b7 22:03 UTC","period_start":"2026-10-04T18:03:01+00:00","prompt_eval_count":2246,"summary":"- glm-5.3 is the strongest model with an average throughput of 123.05 token/s (peak 178.26 token/s), while nemotron-3-ultra is the weakest at 29.47 token/s average, never exceeding 50.9 token/s.\n- deepseek-v4.1-flash shows the most operationally significant volatility, swinging between 19.01 and 194.62 token/s with a coefficient of variation of 53.6%; glm-5.2 is similarly erratic at 59.2%, dipping to 12.85 token/s at 21:00 UTC.\n- No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- glm-5.3 is the strongest model with an average throughput of 123.05 token/s (peak 178.26 token/s), while nemotron-3-ultra is the weakest at 29.47 token/s average, never exceeding 50.9 token/s.\n- deepseek-v4.1-flash shows the most operationally significant volatility, swinging between 19.01 and 194.62 token/s with a coefficient of variation of 53.6%; glm-5.2 is similarly erratic at 59.2%, dipping to 12.85 token/s at 21:00 UTC.\n- No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["glm-5.3 is the strongest model with an average throughput of 123.05 token/s (peak 178.26 token/s), while nemotron-3-ultra is the weakest at 29.47 token/s average, never exceeding 50.9 token/s.","deepseek-v4.1-flash shows the most operationally significant volatility, swinging between 19.01 and 194.62 token/s with a coefficient of variation of 53.6%; glm-5.2 is similarly erratic at 59.2%, dipping to 12.85 token/s at 21:00 UTC.","No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- glm-5.3 \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 123.05 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 178.26 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 29.47 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 50.9 \u8bcd\u5143/\u79d2\u3002\n- deepseek-v4.1-flash \u8868\u73b0\u51fa\u5bf9\u8fd0\u8425\u5f71\u54cd\u6700\u5927\u7684\u6ce2\u52a8\u6027\uff0c\u5728 19.01 \u81f3 194.62 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6446\u52a8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 53.6%\uff1bglm-5.2 \u540c\u6837\u4e0d\u7a33\u5b9a\uff0c\u8fbe 59.2%\uff0c\u5728 21:00 UTC \u65f6\u964d\u81f3 12.85 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8 8 \u4e2a\u6a21\u578b\u5747\u4ea4\u4ed8\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u4ea7\u751f 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":3984234857,"translated_at":"2026-10-04T22:03:07.453144+00:00","translation_eval_count":192,"translation_model":"glm-5.3","translation_prompt_eval_count":340,"translation_total_duration_ns":1877873034,"translation_wall_seconds":2.027,"valid_point_count":96,"wall_seconds":4.151},{"coverage_pct":100.0,"eval_count":642,"generated_at":"2026-10-04T21:03:05.841985+00:00","generated_label":"Oct 04, 2026 \u00b7 21:03 UTC","id":1135,"model":"glm-5.3","period_end":"2026-10-04T21:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 17:03 UTC to Oct 04, 2026 \u00b7 21:03 UTC","period_start":"2026-10-04T17:03:01+00:00","prompt_eval_count":2245,"summary":"- glm-5.3 is the strongest model at 130.48 token/s average throughput, narrowly ahead of deepseek-v4.1-flash at 128.91 token/s; nemotron-3-ultra is the weakest at 24.05 token/s average, never exceeding 41.31 token/s.\n- deepseek-v4.1-flash shows the most operationally significant volatility, swinging from 194.62 token/s at 18:40 to 19.01 token/s at 19:40, with a 47.7% coefficient of variation and a -47.3% trend; glm-5.2 is similarly unstable at 73.3% CV, ending at 12.85 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_en":"- glm-5.3 is the strongest model at 130.48 token/s average throughput, narrowly ahead of deepseek-v4.1-flash at 128.91 token/s; nemotron-3-ultra is the weakest at 24.05 token/s average, never exceeding 41.31 token/s.\n- deepseek-v4.1-flash shows the most operationally significant volatility, swinging from 194.62 token/s at 18:40 to 19.01 token/s at 19:40, with a 47.7% coefficient of variation and a -47.3% trend; glm-5.2 is similarly unstable at 73.3% CV, ending at 12.85 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_items":["glm-5.3 is the strongest model at 130.48 token/s average throughput, narrowly ahead of deepseek-v4.1-flash at 128.91 token/s; nemotron-3-ultra is the weakest at 24.05 token/s average, never exceeding 41.31 token/s.","deepseek-v4.1-flash shows the most operationally significant volatility, swinging from 194.62 token/s at 18:40 to 19.01 token/s at 19:40, with a 47.7% coefficient of variation and a -47.3% trend; glm-5.2 is similarly unstable at 73.3% CV, ending at 12.85 token/s.","No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window."],"summary_zh":"- glm-5.3 \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 130.48 \u8bcd\u5143/\u79d2\uff0c\u7565\u5fae\u9886\u5148\u4e8e deepseek-v4.1-flash \u7684 128.91 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 24.05 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 41.31 \u8bcd\u5143/\u79d2\u3002\n- deepseek-v4.1-flash \u8868\u73b0\u51fa\u5bf9\u8fd0\u8425\u5f71\u54cd\u6700\u663e\u8457\u7684\u6ce2\u52a8\u6027\uff0c\u4ece 18:40 \u7684 194.62 \u8bcd\u5143/\u79d2\u9aa4\u964d\u81f3 19:40 \u7684 19.01 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 47.7%\uff0c\u8d8b\u52bf\u4e3a -47.3%\uff1bglm-5.2 \u540c\u6837\u4e0d\u7a33\u5b9a\uff0cCV \u4e3a 73.3%\uff0c\u6700\u7ec8\u6536\u4e8e 12.85 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u8986\u76d6\u7387\u8fbe\u5230 100.0%\u3002","total_duration_ns":3081744391,"translated_at":"2026-10-04T21:03:05.841985+00:00","translation_eval_count":223,"translation_model":"glm-5.3","translation_prompt_eval_count":362,"translation_total_duration_ns":1027230940,"translation_wall_seconds":1.176,"valid_point_count":96,"wall_seconds":3.253},{"coverage_pct":100.0,"eval_count":781,"generated_at":"2026-10-04T20:03:07.093000+00:00","generated_label":"Oct 04, 2026 \u00b7 20:03 UTC","id":1134,"model":"glm-5.3","period_end":"2026-10-04T20:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 16:03 UTC to Oct 04, 2026 \u00b7 20:03 UTC","period_start":"2026-10-04T16:03:01+00:00","prompt_eval_count":2245,"summary":"- glm-5.3 delivered the highest average throughput at 131.79 token/s, peaking at 178.26 token/s at 20:00 UTC, while nemotron-3-ultra was weakest at 24.09 token/s average and never exceeded 41.31 token/s.\n- deepseek-v4.1-flash showed the most operationally significant volatility, swinging from 12.98 to 198.28 token/s with a 61.1% coefficient of variation, including a drop from 160.86 token/s at 19:00 to 38.51 token/s at 19:20 before recovering to 185.99 token/s.\n- No missing-data limitation applies: all 8 models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- glm-5.3 delivered the highest average throughput at 131.79 token/s, peaking at 178.26 token/s at 20:00 UTC, while nemotron-3-ultra was weakest at 24.09 token/s average and never exceeded 41.31 token/s.\n- deepseek-v4.1-flash showed the most operationally significant volatility, swinging from 12.98 to 198.28 token/s with a 61.1% coefficient of variation, including a drop from 160.86 token/s at 19:00 to 38.51 token/s at 19:20 before recovering to 185.99 token/s.\n- No missing-data limitation applies: all 8 models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["glm-5.3 delivered the highest average throughput at 131.79 token/s, peaking at 178.26 token/s at 20:00 UTC, while nemotron-3-ultra was weakest at 24.09 token/s average and never exceeded 41.31 token/s.","deepseek-v4.1-flash showed the most operationally significant volatility, swinging from 12.98 to 198.28 token/s with a 61.1% coefficient of variation, including a drop from 160.86 token/s at 19:00 to 38.51 token/s at 19:20 before recovering to 185.99 token/s.","No missing-data limitation applies: all 8 models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- glm-5.3 \u7684\u5e73\u5747\u541e\u5410\u91cf\u6700\u9ad8\uff0c\u8fbe\u5230 131.79 \u8bcd\u5143/\u79d2\uff0c\u5e76\u5728 UTC 20:00 \u8fbe\u5230 178.26 \u8bcd\u5143/\u79d2\u7684\u5cf0\u503c\uff0c\u800c nemotron-3-ultra \u8868\u73b0\u6700\u5f31\uff0c\u5e73\u5747\u4ec5\u4e3a 24.09 \u8bcd\u5143/\u79d2\uff0c\u4e14\u4ece\u672a\u8d85\u8fc7 41.31 \u8bcd\u5143/\u79d2\u3002\n- deepseek-v4.1-flash \u8868\u73b0\u51fa\u5bf9\u8fd0\u884c\u5f71\u54cd\u6700\u4e3a\u663e\u8457\u7684\u6ce2\u52a8\u6027\uff0c\u5728 12.98 \u81f3 198.28 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u5927\u5e45\u6446\u52a8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 61.1%\uff0c\u5176\u4e2d\u5305\u62ec\u4ece 19:00 \u7684 160.86 \u8bcd\u5143/\u79d2\u8dcc\u81f3 19:20 \u7684 38.51 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u56de\u5347\u81f3 185.99 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8 8 \u4e2a\u6a21\u578b\u5747\u8bb0\u5f55\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u4ea7\u751f 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u8986\u76d6\u7387\u8fbe\u5230 100.0%\u3002","total_duration_ns":3867769550,"translated_at":"2026-10-04T20:03:07.093000+00:00","translation_eval_count":228,"translation_model":"glm-5.3","translation_prompt_eval_count":351,"translation_total_duration_ns":1062162075,"translation_wall_seconds":1.209,"valid_point_count":96,"wall_seconds":4.036},{"coverage_pct":100.0,"eval_count":309,"generated_at":"2026-10-04T19:03:07.775810+00:00","generated_label":"Oct 04, 2026 \u00b7 19:03 UTC","id":1133,"model":"glm-5.3","period_end":"2026-10-04T19:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 15:03 UTC to Oct 04, 2026 \u00b7 19:03 UTC","period_start":"2026-10-04T15:03:01+00:00","prompt_eval_count":2246,"summary":"- deepseek-v4.1-flash is the strongest model by average throughput at 136.12 token/s, while nemotron-3-ultra is the weakest at 21.19 token/s, roughly a 6.4x gap between the two.\n- The most operationally significant volatility is deepseek-v4.1-flash, which swung from 209.88 token/s at 15:40 down to 12.98 token/s at 16:40, with a coefficient of variation of 46.9%; deepseek-v4-pro also climbed from a 15.9 token/s low to 115.41 token/s.\n- No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model by average throughput at 136.12 token/s, while nemotron-3-ultra is the weakest at 21.19 token/s, roughly a 6.4x gap between the two.\n- The most operationally significant volatility is deepseek-v4.1-flash, which swung from 209.88 token/s at 15:40 down to 12.98 token/s at 16:40, with a coefficient of variation of 46.9%; deepseek-v4-pro also climbed from a 15.9 token/s low to 115.41 token/s.\n- No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model by average throughput at 136.12 token/s, while nemotron-3-ultra is the weakest at 21.19 token/s, roughly a 6.4x gap between the two.","The most operationally significant volatility is deepseek-v4.1-flash, which swung from 209.88 token/s at 15:40 down to 12.98 token/s at 16:40, with a coefficient of variation of 46.9%; deepseek-v4-pro also climbed from a 15.9 token/s low to 115.41 token/s.","No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u8fbe\u5230 136.12 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4ec5\u4e3a 21.19 \u8bcd\u5143/\u79d2\uff0c\u4e24\u8005\u4e4b\u95f4\u5927\u7ea6\u6709 6.4 \u500d\u7684\u5dee\u8ddd\u3002\n- \u8fd0\u8425\u5c42\u9762\u6700\u663e\u8457\u7684\u6ce2\u52a8\u6765\u81ea deepseek-v4.1-flash\uff0c\u5176\u541e\u5410\u91cf\u4ece 15:40 \u7684 209.88 \u8bcd\u5143/\u79d2\u9aa4\u964d\u81f3 16:40 \u7684 12.98 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 46.9%\uff1bdeepseek-v4-pro \u4e5f\u4ece 15.9 \u8bcd\u5143/\u79d2\u7684\u4f4e\u70b9\u6500\u5347\u81f3 115.41 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u8bb0\u5f55\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u5171\u6709 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":3082280272,"translated_at":"2026-10-04T19:03:07.775810+00:00","translation_eval_count":194,"translation_model":"glm-5.3","translation_prompt_eval_count":343,"translation_total_duration_ns":2670223026,"translation_wall_seconds":2.826,"valid_point_count":96,"wall_seconds":3.266},{"coverage_pct":100.0,"eval_count":546,"generated_at":"2026-10-04T18:03:05.676056+00:00","generated_label":"Oct 04, 2026 \u00b7 18:03 UTC","id":1132,"model":"glm-5.3","period_end":"2026-10-04T18:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 14:03 UTC to Oct 04, 2026 \u00b7 18:03 UTC","period_start":"2026-10-04T14:03:01+00:00","prompt_eval_count":2247,"summary":"- deepseek-v4.1-flash is the strongest model by average throughput at 145.0 token/s, with a peak of 232.5 token/s; nemotron-3-ultra is the weakest at 17.82 token/s average, never exceeding 30.41 token/s.\n- The most operationally significant volatility is deepseek-v4.1-flash, which fell from 232.5 token/s at 14:40 to 12.98 token/s at 16:40 before recovering to 198.28 token/s at 17:40, a coefficient of variation of 49.2%; glm-5.2 similarly swung between 101.73 and 15.97 token/s.\n- No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model by average throughput at 145.0 token/s, with a peak of 232.5 token/s; nemotron-3-ultra is the weakest at 17.82 token/s average, never exceeding 30.41 token/s.\n- The most operationally significant volatility is deepseek-v4.1-flash, which fell from 232.5 token/s at 14:40 to 12.98 token/s at 16:40 before recovering to 198.28 token/s at 17:40, a coefficient of variation of 49.2%; glm-5.2 similarly swung between 101.73 and 15.97 token/s.\n- No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model by average throughput at 145.0 token/s, with a peak of 232.5 token/s; nemotron-3-ultra is the weakest at 17.82 token/s average, never exceeding 30.41 token/s.","The most operationally significant volatility is deepseek-v4.1-flash, which fell from 232.5 token/s at 14:40 to 12.98 token/s at 16:40 before recovering to 198.28 token/s at 17:40, a coefficient of variation of 49.2%; glm-5.2 similarly swung between 101.73 and 15.97 token/s.","No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u8fbe\u5230 145.0 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u4e3a 232.5 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4ec5 17.82 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 30.41 \u8bcd\u5143/\u79d2\u3002\n- \u8fd0\u8425\u5c42\u9762\u6700\u663e\u8457\u7684\u6ce2\u52a8\u51fa\u73b0\u5728 deepseek-v4.1-flash\uff0c\u5176\u541e\u5410\u91cf\u4ece 14:40 \u7684 232.5 \u8bcd\u5143/\u79d2\u964d\u81f3 16:40 \u7684 12.98 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u5728 17:40 \u6062\u590d\u81f3 198.28 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 49.2%\uff1bglm-5.2 \u540c\u6837\u5728 101.73 \u548c 15.97 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u5927\u5e45\u6ce2\u52a8\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u8bb0\u5f55\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u5171\u6709 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":2473379811,"translated_at":"2026-10-04T18:03:05.676056+00:00","translation_eval_count":234,"translation_model":"glm-5.3","translation_prompt_eval_count":359,"translation_total_duration_ns":1456097433,"translation_wall_seconds":1.608,"valid_point_count":96,"wall_seconds":2.638},{"coverage_pct":100.0,"eval_count":962,"generated_at":"2026-10-04T17:03:07.421403+00:00","generated_label":"Oct 04, 2026 \u00b7 17:03 UTC","id":1131,"model":"glm-5.3","period_end":"2026-10-04T17:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 13:03 UTC to Oct 04, 2026 \u00b7 17:03 UTC","period_start":"2026-10-04T13:03:01+00:00","prompt_eval_count":2249,"summary":"- deepseek-v4.1-flash is the strongest model at 137.49 token/s average throughput, while nemotron-3-ultra is the weakest at 20.03 token/s average, roughly one-seventh of the leader's pace.\n- The most operationally significant movement is deepseek-v4.1-flash's collapse: after peaking at 243.68 token/s at 13:40 UTC, it fell to 12.98 token/s at 16:40, a -39.7% trend with 60.4% coefficient of variation, indicating severe instability in the top performer.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 137.49 token/s average throughput, while nemotron-3-ultra is the weakest at 20.03 token/s average, roughly one-seventh of the leader's pace.\n- The most operationally significant movement is deepseek-v4.1-flash's collapse: after peaking at 243.68 token/s at 13:40 UTC, it fell to 12.98 token/s at 16:40, a -39.7% trend with 60.4% coefficient of variation, indicating severe instability in the top performer.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 137.49 token/s average throughput, while nemotron-3-ultra is the weakest at 20.03 token/s average, roughly one-seventh of the leader's pace.","The most operationally significant movement is deepseek-v4.1-flash's collapse: after peaking at 243.68 token/s at 13:40 UTC, it fell to 12.98 token/s at 16:40, a -39.7% trend with 60.4% coefficient of variation, indicating severe instability in the top performer.","No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 137.49 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u662f\u6700\u5f31\u6a21\u578b\uff0c\u5e73\u5747\u4e3a 20.03 \u8bcd\u5143/\u79d2\uff0c\u7ea6\u4e3a\u9886\u5148\u8005\u901f\u5ea6\u7684\u4e03\u5206\u4e4b\u4e00\u3002\n- \u8fd0\u8425\u5c42\u9762\u6700\u663e\u8457\u7684\u53d8\u5316\u662f deepseek-v4.1-flash \u7684\u5d29\u6e83\uff1a\u5728 13:40 UTC \u8fbe\u5230 243.68 \u8bcd\u5143/\u79d2\u7684\u5cf0\u503c\u540e\uff0c\u4e8e 16:40 \u8dcc\u81f3 12.98 \u8bcd\u5143/\u79d2\uff0c\u8d8b\u52bf\u4e3a -39.7%\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 60.4%\uff0c\u8868\u660e\u8868\u73b0\u6700\u4f73\u7684\u6a21\u578b\u5b58\u5728\u4e25\u91cd\u7684\u4e0d\u7a33\u5b9a\u6027\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0cvalid_point_count \u4e3a 96\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":4034387623,"translated_at":"2026-10-04T17:03:07.421403+00:00","translation_eval_count":214,"translation_model":"glm-5.3","translation_prompt_eval_count":340,"translation_total_duration_ns":1059372721,"translation_wall_seconds":1.21,"valid_point_count":96,"wall_seconds":4.382},{"coverage_pct":100.0,"eval_count":453,"generated_at":"2026-10-04T16:03:18.620185+00:00","generated_label":"Oct 04, 2026 \u00b7 16:03 UTC","id":1130,"model":"glm-5.3","period_end":"2026-10-04T16:03:02+00:00","period_label":"Oct 04, 2026 \u00b7 12:03 UTC to Oct 04, 2026 \u00b7 16:03 UTC","period_start":"2026-10-04T12:03:02+00:00","prompt_eval_count":2249,"summary":"- deepseek-v4.1-flash is the strongest model at 150.03 token/s average throughput; nemotron-3-ultra is the weakest at 19.73 token/s average, roughly 7.6 times lower.\n- deepseek-v4.1-flash shows the most operationally significant volatility, ranging from 20.22 to 243.68 token/s with a 51.5% coefficient of variation and a +63.2% trend, while minimax-m3 declined 42.4% to a 42.08 token/s latest reading.\n- No missing-data limitation applies: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 150.03 token/s average throughput; nemotron-3-ultra is the weakest at 19.73 token/s average, roughly 7.6 times lower.\n- deepseek-v4.1-flash shows the most operationally significant volatility, ranging from 20.22 to 243.68 token/s with a 51.5% coefficient of variation and a +63.2% trend, while minimax-m3 declined 42.4% to a 42.08 token/s latest reading.\n- No missing-data limitation applies: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 150.03 token/s average throughput; nemotron-3-ultra is the weakest at 19.73 token/s average, roughly 7.6 times lower.","deepseek-v4.1-flash shows the most operationally significant volatility, ranging from 20.22 to 243.68 token/s with a 51.5% coefficient of variation and a +63.2% trend, while minimax-m3 declined 42.4% to a 42.08 token/s latest reading.","No missing-data limitation applies: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 150.03 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 19.73 \u8bcd\u5143/\u79d2\uff0c\u4f4e\u4e86\u7ea6 7.6 \u500d\u3002\n- deepseek-v4.1-flash \u7684\u6ce2\u52a8\u6027\u6700\u5177\u8fd0\u8425\u610f\u4e49\uff0c\u8303\u56f4\u4ece 20.22 \u5230 243.68 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 51.5%\uff0c\u8d8b\u52bf\u4e3a +63.2%\uff0c\u800c minimax-m3 \u4e0b\u964d\u4e86 42.4%\uff0c\u6700\u65b0\u8bfb\u6570\u4e3a 42.08 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62a5\u544a 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\u300196 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":3644858794,"translated_at":"2026-10-04T16:03:18.620185+00:00","translation_eval_count":169,"translation_model":"glm-5.3","translation_prompt_eval_count":329,"translation_total_duration_ns":12367455608,"translation_wall_seconds":12.515,"valid_point_count":96,"wall_seconds":4.013},{"coverage_pct":100.0,"eval_count":834,"generated_at":"2026-10-04T15:03:06.614620+00:00","generated_label":"Oct 04, 2026 \u00b7 15:03 UTC","id":1129,"model":"glm-5.3","period_end":"2026-10-04T15:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 11:03 UTC to Oct 04, 2026 \u00b7 15:03 UTC","period_start":"2026-10-04T11:03:01+00:00","prompt_eval_count":2248,"summary":"- deepseek-v4.1-flash is the strongest model at 135.07 token/s average throughput, while nemotron-3-ultra is the weakest at 24.51 token/s average.\n- The most operationally significant volatility is deepseek-v4.1-flash, whose throughput swings between 20.22 and 243.68 token/s (coefficient of variation 57.9%), while gemma4:31b shows the steepest decline, falling 40.5% from a 184.32 token/s peak to roughly 68-97 token/s late in the window.\n- No missing-data limitation exists: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage; the only limitation is the short four-hour window with twelve observations per model.","summary_en":"- deepseek-v4.1-flash is the strongest model at 135.07 token/s average throughput, while nemotron-3-ultra is the weakest at 24.51 token/s average.\n- The most operationally significant volatility is deepseek-v4.1-flash, whose throughput swings between 20.22 and 243.68 token/s (coefficient of variation 57.9%), while gemma4:31b shows the steepest decline, falling 40.5% from a 184.32 token/s peak to roughly 68-97 token/s late in the window.\n- No missing-data limitation exists: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage; the only limitation is the short four-hour window with twelve observations per model.","summary_items":["deepseek-v4.1-flash is the strongest model at 135.07 token/s average throughput, while nemotron-3-ultra is the weakest at 24.51 token/s average.","The most operationally significant volatility is deepseek-v4.1-flash, whose throughput swings between 20.22 and 243.68 token/s (coefficient of variation 57.9%), while gemma4:31b shows the steepest decline, falling 40.5% from a 184.32 token/s peak to roughly 68-97 token/s late in the window.","No missing-data limitation exists: all eight models report 12 of 12 samples, 96 valid points, and 100.0% coverage; the only limitation is the short four-hour window with twelve observations per model."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 135.07 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 24.51 \u8bcd\u5143/\u79d2\u3002\n- \u8fd0\u8425\u5c42\u9762\u6700\u663e\u8457\u7684\u6ce2\u52a8\u6027\u6765\u81ea deepseek-v4.1-flash\uff0c\u5176\u541e\u5410\u91cf\u5728 20.22 \u81f3 243.68 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\uff08\u53d8\u5f02\u7cfb\u6570 57.9%\uff09\uff0c\u800c gemma4:31b \u5448\u73b0\u6700\u9661\u5ced\u7684\u4e0b\u964d\uff0c\u4ece 184.32 \u8bcd\u5143/\u79d2\u7684\u5cf0\u503c\u4e0b\u964d 40.5%\uff0c\u5728\u65f6\u95f4\u7a97\u53e3\u540e\u671f\u964d\u81f3\u7ea6 68-97 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62a5\u544a 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\u300196 \u4e2a\u6709\u6548\u6570\u636e\u70b9\u4ee5\u53ca 100.0% \u7684\u8986\u76d6\u7387\uff1b\u552f\u4e00\u7684\u9650\u5236\u662f\u65f6\u95f4\u7a97\u53e3\u8f83\u77ed\uff0c\u4ec5\u4e3a\u56db\u5c0f\u65f6\uff0c\u6bcf\u4e2a\u6a21\u578b\u6709\u5341\u4e8c\u6b21\u89c2\u6d4b\u3002","total_duration_ns":3870043141,"translated_at":"2026-10-04T15:03:06.614620+00:00","translation_eval_count":210,"translation_model":"glm-5.3","translation_prompt_eval_count":345,"translation_total_duration_ns":1000198567,"translation_wall_seconds":1.152,"valid_point_count":96,"wall_seconds":4.233},{"coverage_pct":100.0,"eval_count":547,"generated_at":"2026-10-04T14:03:05.527511+00:00","generated_label":"Oct 04, 2026 \u00b7 14:03 UTC","id":1128,"model":"glm-5.3","period_end":"2026-10-04T14:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 10:03 UTC to Oct 04, 2026 \u00b7 14:03 UTC","period_start":"2026-10-04T10:03:01+00:00","prompt_eval_count":2246,"summary":"- Strongest average throughput was gemma4:31b at 135.0 token/s, peaking at 184.32 token/s at 11:40, while nemotron-3-ultra was weakest at 28.03 token/s average and never exceeded 57.64 token/s.\n- The most operationally significant volatility came from deepseek-v4.1-flash, which ranged from 243.68 token/s at 13:40 down to 20.22 token/s at 14:00 with a 56.2% coefficient of variation; deepseek-v4-pro similarly swung between 11.76 and 114.54 token/s.\n- No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- Strongest average throughput was gemma4:31b at 135.0 token/s, peaking at 184.32 token/s at 11:40, while nemotron-3-ultra was weakest at 28.03 token/s average and never exceeded 57.64 token/s.\n- The most operationally significant volatility came from deepseek-v4.1-flash, which ranged from 243.68 token/s at 13:40 down to 20.22 token/s at 14:00 with a 56.2% coefficient of variation; deepseek-v4-pro similarly swung between 11.76 and 114.54 token/s.\n- No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["Strongest average throughput was gemma4:31b at 135.0 token/s, peaking at 184.32 token/s at 11:40, while nemotron-3-ultra was weakest at 28.03 token/s average and never exceeded 57.64 token/s.","The most operationally significant volatility came from deepseek-v4.1-flash, which ranged from 243.68 token/s at 13:40 down to 20.22 token/s at 14:00 with a 56.2% coefficient of variation; deepseek-v4-pro similarly swung between 11.76 and 114.54 token/s.","No missing-data limitation applies: all eight models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- \u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u662f gemma4:31b\uff0c\u8fbe\u5230 135.0 \u8bcd\u5143/\u79d2\uff0c\u5728 11:40 \u8fbe\u5230\u5cf0\u503c 184.32 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 28.03 \u8bcd\u5143/\u79d2\uff0c\u4e14\u4ece\u672a\u8d85\u8fc7 57.64 \u8bcd\u5143/\u79d2\u3002\n- \u8fd0\u8425\u5c42\u9762\u6700\u663e\u8457\u7684\u6ce2\u52a8\u6765\u81ea deepseek-v4.1-flash\uff0c\u5176\u541e\u5410\u91cf\u5728 13:40 \u4e3a 243.68 \u8bcd\u5143/\u79d2\uff0c\u5230 14:00 \u964d\u81f3 20.22 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 56.2%\uff1bdeepseek-v4-pro \u540c\u6837\u5728 11.76 \u81f3 114.54 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u5927\u5e45\u6ce2\u52a8\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u4ea4\u4ed8\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u4ea7\u751f 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":2403519903,"translated_at":"2026-10-04T14:03:05.527511+00:00","translation_eval_count":218,"translation_model":"glm-5.3","translation_prompt_eval_count":349,"translation_total_duration_ns":1091462801,"translation_wall_seconds":1.235,"valid_point_count":96,"wall_seconds":2.571},{"coverage_pct":100.0,"eval_count":562,"generated_at":"2026-10-04T13:03:05.925009+00:00","generated_label":"Oct 04, 2026 \u00b7 13:03 UTC","id":1127,"model":"glm-5.3","period_end":"2026-10-04T13:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 09:03 UTC to Oct 04, 2026 \u00b7 13:03 UTC","period_start":"2026-10-04T09:03:01+00:00","prompt_eval_count":2244,"summary":"- gemma4:31b delivered the highest average throughput at 135.78 token/s, while nemotron-3-ultra was weakest at 27.62 token/s; deepseek-v4.1-flash and glm-5.3 followed closely at 132.95 and 129.78 token/s.\n- deepseek-v4.1-flash showed the sharpest swing, peaking at 210.67 token/s at 10:00 before falling to 37.08 token/s at 12:40, with a coefficient of variation of 44.3% and a trend of -41.1%; glm-5.3 also swung between 182.87 and 65.78 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, though the window covers only four hours.","summary_en":"- gemma4:31b delivered the highest average throughput at 135.78 token/s, while nemotron-3-ultra was weakest at 27.62 token/s; deepseek-v4.1-flash and glm-5.3 followed closely at 132.95 and 129.78 token/s.\n- deepseek-v4.1-flash showed the sharpest swing, peaking at 210.67 token/s at 10:00 before falling to 37.08 token/s at 12:40, with a coefficient of variation of 44.3% and a trend of -41.1%; glm-5.3 also swung between 182.87 and 65.78 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, though the window covers only four hours.","summary_items":["gemma4:31b delivered the highest average throughput at 135.78 token/s, while nemotron-3-ultra was weakest at 27.62 token/s; deepseek-v4.1-flash and glm-5.3 followed closely at 132.95 and 129.78 token/s.","deepseek-v4.1-flash showed the sharpest swing, peaking at 210.67 token/s at 10:00 before falling to 37.08 token/s at 12:40, with a coefficient of variation of 44.3% and a trend of -41.1%; glm-5.3 also swung between 182.87 and 65.78 token/s.","No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, though the window covers only four hours."],"summary_zh":"- gemma4:31b \u7684\u5e73\u5747\u541e\u5410\u91cf\u6700\u9ad8\uff0c\u8fbe\u5230 135.78 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u4ec5\u4e3a 27.62 \u8bcd\u5143/\u79d2\uff1bdeepseek-v4.1-flash \u548c glm-5.3 \u7d27\u968f\u5176\u540e\uff0c\u5206\u522b\u4e3a 132.95 \u548c 129.78 \u8bcd\u5143/\u79d2\u3002\n- deepseek-v4.1-flash \u7684\u6ce2\u52a8\u6700\u4e3a\u5267\u70c8\uff0c\u5728 10:00 \u8fbe\u5230\u5cf0\u503c 210.67 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u5728 12:40 \u964d\u81f3 37.08 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 44.3%\uff0c\u8d8b\u52bf\u4e3a -41.1%\uff1bglm-5.3 \u4e5f\u5728 182.87 \u548c 65.78 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\u3002\n- \u4e0d\u5b58\u5728\u7f3a\u5931\u6570\u636e\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u4e0d\u8fc7\u8be5\u65f6\u95f4\u7a97\u53e3\u4ec5\u8986\u76d6\u56db\u4e2a\u5c0f\u65f6\u3002","total_duration_ns":2579078900,"translated_at":"2026-10-04T13:03:05.925009+00:00","translation_eval_count":223,"translation_model":"glm-5.3","translation_prompt_eval_count":361,"translation_total_duration_ns":1131532628,"translation_wall_seconds":1.278,"valid_point_count":96,"wall_seconds":2.745},{"coverage_pct":100.0,"eval_count":886,"generated_at":"2026-10-04T12:03:06.986819+00:00","generated_label":"Oct 04, 2026 \u00b7 12:03 UTC","id":1126,"model":"glm-5.3","period_end":"2026-10-04T12:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 08:03 UTC to Oct 04, 2026 \u00b7 12:03 UTC","period_start":"2026-10-04T08:03:01+00:00","prompt_eval_count":2247,"summary":"- deepseek-v4.1-flash posted the highest average throughput at 140.28 token/s, peaking at 210.67 token/s, while nemotron-3-ultra was weakest at 25.28 token/s average and never exceeded 45.36 token/s.\n- glm-5.3 showed the sharpest volatility, ranging from 50.69 to 182.87 token/s with a 47.2% coefficient of variation, including a fall from 181.17 token/s at 10:20 to 72.23 token/s at 10:40; deepseek-v4.1-flash also dropped from 142.01 to 37.34 token/s at 12:00.\n- No missing-data limitation applies: all 8 models recorded 12 of 12 expected samples, totaling 96 valid points with 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash posted the highest average throughput at 140.28 token/s, peaking at 210.67 token/s, while nemotron-3-ultra was weakest at 25.28 token/s average and never exceeded 45.36 token/s.\n- glm-5.3 showed the sharpest volatility, ranging from 50.69 to 182.87 token/s with a 47.2% coefficient of variation, including a fall from 181.17 token/s at 10:20 to 72.23 token/s at 10:40; deepseek-v4.1-flash also dropped from 142.01 to 37.34 token/s at 12:00.\n- No missing-data limitation applies: all 8 models recorded 12 of 12 expected samples, totaling 96 valid points with 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash posted the highest average throughput at 140.28 token/s, peaking at 210.67 token/s, while nemotron-3-ultra was weakest at 25.28 token/s average and never exceeded 45.36 token/s.","glm-5.3 showed the sharpest volatility, ranging from 50.69 to 182.87 token/s with a 47.2% coefficient of variation, including a fall from 181.17 token/s at 10:20 to 72.23 token/s at 10:40; deepseek-v4.1-flash also dropped from 142.01 to 37.34 token/s at 12:00.","No missing-data limitation applies: all 8 models recorded 12 of 12 expected samples, totaling 96 valid points with 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u7684\u5e73\u5747\u541e\u5410\u91cf\u6700\u9ad8\uff0c\u8fbe 140.28 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u4e3a 210.67 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4ec5 25.28 \u8bcd\u5143/\u79d2\uff0c\u4e14\u4ece\u672a\u8d85\u8fc7 45.36 \u8bcd\u5143/\u79d2\u3002\n- glm-5.3 \u7684\u6ce2\u52a8\u6700\u4e3a\u5267\u70c8\uff0c\u8303\u56f4\u4ece 50.69 \u5230 182.87 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 47.2%\uff0c\u5176\u4e2d\u5305\u62ec\u4ece 10:20 \u7684 181.17 \u8bcd\u5143/\u79d2\u8dcc\u81f3 10:40 \u7684 72.23 \u8bcd\u5143/\u79d2\uff1bdeepseek-v4.1-flash \u4e5f\u5728 12:00 \u4ece 142.01 \u964d\u81f3 37.34 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u7f3a\u5931\u6570\u636e\u7684\u9650\u5236\uff1a\u5168\u90e8 8 \u4e2a\u6a21\u578b\u5747\u8bb0\u5f55\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5171\u8ba1 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u8986\u76d6\u7387\u8fbe\u5230 100.0%\u3002","total_duration_ns":3672284725,"translated_at":"2026-10-04T12:03:06.986819+00:00","translation_eval_count":214,"translation_model":"glm-5.3","translation_prompt_eval_count":361,"translation_total_duration_ns":1315480292,"translation_wall_seconds":1.457,"valid_point_count":96,"wall_seconds":3.836},{"coverage_pct":100.0,"eval_count":518,"generated_at":"2026-10-04T11:03:04.797801+00:00","generated_label":"Oct 04, 2026 \u00b7 11:03 UTC","id":1125,"model":"glm-5.3","period_end":"2026-10-04T11:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 07:03 UTC to Oct 04, 2026 \u00b7 11:03 UTC","period_start":"2026-10-04T07:03:01+00:00","prompt_eval_count":2247,"summary":"- glm-5.3-flash is the strongest model at 146.88 token/s average throughput, peaking at 221.2 token/s; nemotron-3-ultra is the weakest at 25.75 token/s average, never exceeding 79.12 token/s.\n- glm-5.3 shows the most operationally significant volatility, swinging between 49.75 and 182.87 token/s with a 46.8% coefficient of variation, while deepseek-v4-pro dropped to 5.64 token/s at 10:00 UTC before recovering to 114.54 token/s by 10:40.\n- No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- glm-5.3-flash is the strongest model at 146.88 token/s average throughput, peaking at 221.2 token/s; nemotron-3-ultra is the weakest at 25.75 token/s average, never exceeding 79.12 token/s.\n- glm-5.3 shows the most operationally significant volatility, swinging between 49.75 and 182.87 token/s with a 46.8% coefficient of variation, while deepseek-v4-pro dropped to 5.64 token/s at 10:00 UTC before recovering to 114.54 token/s by 10:40.\n- No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["glm-5.3-flash is the strongest model at 146.88 token/s average throughput, peaking at 221.2 token/s; nemotron-3-ultra is the weakest at 25.75 token/s average, never exceeding 79.12 token/s.","glm-5.3 shows the most operationally significant volatility, swinging between 49.75 and 182.87 token/s with a 46.8% coefficient of variation, while deepseek-v4-pro dropped to 5.64 token/s at 10:00 UTC before recovering to 114.54 token/s by 10:40.","No missing-data limitation exists: all eight models recorded 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- glm-5.3-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 146.88 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 221.2 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 25.75 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 79.12 \u8bcd\u5143/\u79d2\u3002\n- glm-5.3 \u7684\u6ce2\u52a8\u6027\u6700\u5177\u8fd0\u7ef4\u610f\u4e49\uff0c\u5728 49.75 \u81f3 182.87 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6ce2\u52a8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 46.8%\uff0c\u800c deepseek-v4-pro \u5728 UTC 10:00 \u964d\u81f3 5.64 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u5728 10:40 \u6062\u590d\u81f3 114.54 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u8bb0\u5f55\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u4ea7\u751f 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":2084695638,"translated_at":"2026-10-04T11:03:04.797801+00:00","translation_eval_count":204,"translation_model":"glm-5.3","translation_prompt_eval_count":343,"translation_total_duration_ns":806494314,"translation_wall_seconds":1.031,"valid_point_count":96,"wall_seconds":2.254},{"coverage_pct":100.0,"eval_count":318,"generated_at":"2026-10-04T10:03:04.308756+00:00","generated_label":"Oct 04, 2026 \u00b7 10:03 UTC","id":1124,"model":"glm-5.3","period_end":"2026-10-04T10:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 06:03 UTC to Oct 04, 2026 \u00b7 10:03 UTC","period_start":"2026-10-04T06:03:01+00:00","prompt_eval_count":2249,"summary":"- glm-5.3-flash is the strongest model by average throughput at 147.85 token/s, ahead of deepseek-v4.1-flash (134.79 token/s) and gemma4:31b (133.44 token/s); nemotron-3-ultra is the weakest at 24.36 token/s average, with a maximum of only 79.12 token/s.\n- The most operationally significant volatility is deepseek-v4-pro's collapse to 5.64 token/s at 10:00 after holding 85.61 to 114.36 token/s from 09:00 to 09:40; deepseek-v4.1-flash similarly fell to 24.35 token/s at 08:20 before recovering to 210.67 token/s by 10:00.\n- No missing-data limitation exists: all eight models have 12 of 12 samples, and overall coverage is 100.0% with 96 valid points.","summary_en":"- glm-5.3-flash is the strongest model by average throughput at 147.85 token/s, ahead of deepseek-v4.1-flash (134.79 token/s) and gemma4:31b (133.44 token/s); nemotron-3-ultra is the weakest at 24.36 token/s average, with a maximum of only 79.12 token/s.\n- The most operationally significant volatility is deepseek-v4-pro's collapse to 5.64 token/s at 10:00 after holding 85.61 to 114.36 token/s from 09:00 to 09:40; deepseek-v4.1-flash similarly fell to 24.35 token/s at 08:20 before recovering to 210.67 token/s by 10:00.\n- No missing-data limitation exists: all eight models have 12 of 12 samples, and overall coverage is 100.0% with 96 valid points.","summary_items":["glm-5.3-flash is the strongest model by average throughput at 147.85 token/s, ahead of deepseek-v4.1-flash (134.79 token/s) and gemma4:31b (133.44 token/s); nemotron-3-ultra is the weakest at 24.36 token/s average, with a maximum of only 79.12 token/s.","The most operationally significant volatility is deepseek-v4-pro's collapse to 5.64 token/s at 10:00 after holding 85.61 to 114.36 token/s from 09:00 to 09:40; deepseek-v4.1-flash similarly fell to 24.35 token/s at 08:20 before recovering to 210.67 token/s by 10:00.","No missing-data limitation exists: all eight models have 12 of 12 samples, and overall coverage is 100.0% with 96 valid points."],"summary_zh":"- glm-5.3-flash \u662f\u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u8fbe\u5230 147.85 \u8bcd\u5143/\u79d2\uff0c\u9886\u5148\u4e8e deepseek-v4.1-flash\uff08134.79 \u8bcd\u5143/\u79d2\uff09\u548c gemma4:31b\uff08133.44 \u8bcd\u5143/\u79d2\uff09\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4ec5\u4e3a 24.36 \u8bcd\u5143/\u79d2\uff0c\u6700\u5927\u503c\u4e5f\u4ec5\u6709 79.12 \u8bcd\u5143/\u79d2\u3002\n- \u8fd0\u8425\u5c42\u9762\u6700\u663e\u8457\u7684\u6ce2\u52a8\u662f deepseek-v4-pro \u5728 09:00 \u81f3 09:40 \u671f\u95f4\u4fdd\u6301 85.61 \u81f3 114.36 \u8bcd\u5143/\u79d2\u540e\uff0c\u4e8e 10:00 \u9aa4\u964d\u81f3 5.64 \u8bcd\u5143/\u79d2\uff1bdeepseek-v4.1-flash \u540c\u6837\u5728 08:20 \u8dcc\u81f3 24.35 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u5728 10:00 \u524d\u6062\u590d\u81f3 210.67 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u6574\u4f53\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\u3002","total_duration_ns":1545075544,"translated_at":"2026-10-04T10:03:04.308756+00:00","translation_eval_count":247,"translation_model":"glm-5.3","translation_prompt_eval_count":376,"translation_total_duration_ns":1144580214,"translation_wall_seconds":1.295,"valid_point_count":96,"wall_seconds":1.714},{"coverage_pct":100.0,"eval_count":573,"generated_at":"2026-10-04T09:03:09.705769+00:00","generated_label":"Oct 04, 2026 \u00b7 09:03 UTC","id":1123,"model":"glm-5.3","period_end":"2026-10-04T09:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 05:03 UTC to Oct 04, 2026 \u00b7 09:03 UTC","period_start":"2026-10-04T05:03:01+00:00","prompt_eval_count":2249,"summary":"- gemma4:31b is the strongest model by average throughput at 142.88 token/s, narrowly ahead of glm-5.3-flash at 137.72 token/s; nemotron-3-ultra is the weakest at 24.25 token/s average, with a maximum of only 79.12 token/s.\n- deepseek-v4.1-flash showed the sharpest deterioration, trending down 41.8% with a fall from 196.28 token/s at 06:00 to 24.35 token/s at 08:20, while glm-5.3-flash improved 57.7% from 5.74 token/s at 05:20 to 221.2 token/s at 09:00; nemotron-3-ultra was the most volatile at 81.5% CV.\n- No missing-data limitation applies: all 96 expected points are valid, coverage is 100.0%, and every model has 12 of 12 samples across the four-hour window.","summary_en":"- gemma4:31b is the strongest model by average throughput at 142.88 token/s, narrowly ahead of glm-5.3-flash at 137.72 token/s; nemotron-3-ultra is the weakest at 24.25 token/s average, with a maximum of only 79.12 token/s.\n- deepseek-v4.1-flash showed the sharpest deterioration, trending down 41.8% with a fall from 196.28 token/s at 06:00 to 24.35 token/s at 08:20, while glm-5.3-flash improved 57.7% from 5.74 token/s at 05:20 to 221.2 token/s at 09:00; nemotron-3-ultra was the most volatile at 81.5% CV.\n- No missing-data limitation applies: all 96 expected points are valid, coverage is 100.0%, and every model has 12 of 12 samples across the four-hour window.","summary_items":["gemma4:31b is the strongest model by average throughput at 142.88 token/s, narrowly ahead of glm-5.3-flash at 137.72 token/s; nemotron-3-ultra is the weakest at 24.25 token/s average, with a maximum of only 79.12 token/s.","deepseek-v4.1-flash showed the sharpest deterioration, trending down 41.8% with a fall from 196.28 token/s at 06:00 to 24.35 token/s at 08:20, while glm-5.3-flash improved 57.7% from 5.74 token/s at 05:20 to 221.2 token/s at 09:00; nemotron-3-ultra was the most volatile at 81.5% CV.","No missing-data limitation applies: all 96 expected points are valid, coverage is 100.0%, and every model has 12 of 12 samples across the four-hour window."],"summary_zh":"- gemma4:31b \u662f\u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u8fbe\u5230 142.88 \u8bcd\u5143/\u79d2\uff0c\u7565\u5fae\u9886\u5148\u4e8e glm-5.3-flash \u7684 137.72 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4ec5\u4e3a 24.25 \u8bcd\u5143/\u79d2\uff0c\u6700\u5927\u503c\u4e5f\u53ea\u6709 79.12 \u8bcd\u5143/\u79d2\u3002\n- deepseek-v4.1-flash \u8868\u73b0\u51fa\u6700\u6025\u5267\u7684\u6076\u5316\uff0c\u5448 41.8% \u7684\u4e0b\u964d\u8d8b\u52bf\uff0c\u4ece 06:00 \u7684 196.28 \u8bcd\u5143/\u79d2\u8dcc\u81f3 08:20 \u7684 24.35 \u8bcd\u5143/\u79d2\uff0c\u800c glm-5.3-flash \u5219\u6539\u5584\u4e86 57.7%\uff0c\u4ece 05:20 \u7684 5.74 \u8bcd\u5143/\u79d2\u63d0\u5347\u81f3 09:00 \u7684 221.2 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6ce2\u52a8\u6700\u5927\uff0cCV \u4e3a 81.5%\u3002\n- \u4e0d\u5b58\u5728\u7f3a\u5931\u6570\u636e\u7684\u9650\u5236\uff1a\u5168\u90e8 96 \u4e2a\u9884\u671f\u6570\u636e\u70b9\u5747\u6709\u6548\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u4e14\u6bcf\u4e2a\u6a21\u578b\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u90fd\u62e5\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684\u5168\u90e8 12 \u4e2a\u3002","total_duration_ns":5250930998,"translated_at":"2026-10-04T09:03:09.705769+00:00","translation_eval_count":255,"translation_model":"glm-5.3","translation_prompt_eval_count":387,"translation_total_duration_ns":2756487591,"translation_wall_seconds":2.907,"valid_point_count":96,"wall_seconds":5.418},{"coverage_pct":100.0,"eval_count":323,"generated_at":"2026-10-04T08:03:08.810242+00:00","generated_label":"Oct 04, 2026 \u00b7 08:03 UTC","id":1122,"model":"glm-5.3","period_end":"2026-10-04T08:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 04:03 UTC to Oct 04, 2026 \u00b7 08:03 UTC","period_start":"2026-10-04T04:03:01+00:00","prompt_eval_count":2250,"summary":"- gemma4:31b is the strongest model by average throughput at 152.3 token/s (p95 175.61 token/s), while nemotron-3-ultra is the weakest at 28.44 token/s average, peaking at only 79.12 token/s.\n- The most operationally significant movement is glm-5.3-flash, which climbed from 5.15 token/s at 04:20 to 233.42 token/s at 06:20, a 294.8% trend with 77.4% coefficient of variation; deepseek-v4.1-flash also fell to 24.65 token/s at 08:00, its window minimum.\n- No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage, so the four-hour window is complete.","summary_en":"- gemma4:31b is the strongest model by average throughput at 152.3 token/s (p95 175.61 token/s), while nemotron-3-ultra is the weakest at 28.44 token/s average, peaking at only 79.12 token/s.\n- The most operationally significant movement is glm-5.3-flash, which climbed from 5.15 token/s at 04:20 to 233.42 token/s at 06:20, a 294.8% trend with 77.4% coefficient of variation; deepseek-v4.1-flash also fell to 24.65 token/s at 08:00, its window minimum.\n- No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage, so the four-hour window is complete.","summary_items":["gemma4:31b is the strongest model by average throughput at 152.3 token/s (p95 175.61 token/s), while nemotron-3-ultra is the weakest at 28.44 token/s average, peaking at only 79.12 token/s.","The most operationally significant movement is glm-5.3-flash, which climbed from 5.15 token/s at 04:20 to 233.42 token/s at 06:20, a 294.8% trend with 77.4% coefficient of variation; deepseek-v4.1-flash also fell to 24.65 token/s at 08:00, its window minimum.","No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage, so the four-hour window is complete."],"summary_zh":"- gemma4:31b \u662f\u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u8fbe\u5230 152.3 \u8bcd\u5143/\u79d2\uff08p95 \u4e3a 175.61 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4ec5\u4e3a 28.44 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u4e5f\u4ec5\u6709 79.12 \u8bcd\u5143/\u79d2\u3002\n- \u8fd0\u8425\u5c42\u9762\u6700\u663e\u8457\u7684\u53d8\u5316\u6765\u81ea glm-5.3-flash\uff0c\u5176\u4ece 04:20 \u7684 5.15 \u8bcd\u5143/\u79d2\u6500\u5347\u81f3 06:20 \u7684 233.42 \u8bcd\u5143/\u79d2\uff0c\u8d8b\u52bf\u4e3a 294.8%\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 77.4%\uff1bdeepseek-v4.1-flash \u4e5f\u5728 08:00 \u964d\u81f3 24.65 \u8bcd\u5143/\u79d2\uff0c\u4e3a\u5176\u7a97\u53e3\u6700\u4f4e\u503c\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62a5\u544a\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u56e0\u6b64\u8be5\u56db\u5c0f\u65f6\u7a97\u53e3\u662f\u5b8c\u6574\u7684\u3002","total_duration_ns":1535583946,"translated_at":"2026-10-04T08:03:08.810242+00:00","translation_eval_count":222,"translation_model":"glm-5.3","translation_prompt_eval_count":362,"translation_total_duration_ns":5151480315,"translation_wall_seconds":5.31,"valid_point_count":96,"wall_seconds":1.706},{"coverage_pct":100.0,"eval_count":526,"generated_at":"2026-10-04T07:03:07.753995+00:00","generated_label":"Oct 04, 2026 \u00b7 07:03 UTC","id":1121,"model":"glm-5.3","period_end":"2026-10-04T07:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 03:03 UTC to Oct 04, 2026 \u00b7 07:03 UTC","period_start":"2026-10-04T03:03:01+00:00","prompt_eval_count":2250,"summary":"- deepseek-v4.1-flash is the strongest model at 154.31 token/s average, narrowly ahead of gemma4:31b at 151.3 token/s; nemotron-3-ultra is the weakest at 34.64 token/s average, ending at just 3.86 token/s.\n- glm-5.3-flash shows the most operationally significant volatility: it ran near 5-12 token/s for the first seven intervals, then jumped to 233.42 token/s at 06:20, with a coefficient of variation of 120.2%. glm-5.3 also fell from 184.0 token/s at 03:40 to 64.84 token/s at 05:40.\n- No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 154.31 token/s average, narrowly ahead of gemma4:31b at 151.3 token/s; nemotron-3-ultra is the weakest at 34.64 token/s average, ending at just 3.86 token/s.\n- glm-5.3-flash shows the most operationally significant volatility: it ran near 5-12 token/s for the first seven intervals, then jumped to 233.42 token/s at 06:20, with a coefficient of variation of 120.2%. glm-5.3 also fell from 184.0 token/s at 03:40 to 64.84 token/s at 05:40.\n- No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 154.31 token/s average, narrowly ahead of gemma4:31b at 151.3 token/s; nemotron-3-ultra is the weakest at 34.64 token/s average, ending at just 3.86 token/s.","glm-5.3-flash shows the most operationally significant volatility: it ran near 5-12 token/s for the first seven intervals, then jumped to 233.42 token/s at 06:20, with a coefficient of variation of 120.2%. glm-5.3 also fell from 184.0 token/s at 03:40 to 64.84 token/s at 05:40.","No missing-data limitation exists: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0% across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u901f\u5ea6\u4e3a 154.31 \u8bcd\u5143/\u79d2\uff0c\u7565\u5fae\u9886\u5148\u4e8e gemma4:31b \u7684 151.3 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u901f\u5ea6\u4e3a 34.64 \u8bcd\u5143/\u79d2\uff0c\u6700\u7ec8\u4ec5\u4e3a 3.86 \u8bcd\u5143/\u79d2\u3002\n- glm-5.3-flash \u8868\u73b0\u51fa\u5bf9\u8fd0\u884c\u5f71\u54cd\u6700\u4e3a\u663e\u8457\u7684\u6ce2\u52a8\u6027\uff1a\u5b83\u5728\u524d\u4e03\u4e2a\u65f6\u95f4\u95f4\u9694\u5185\u8fd0\u884c\u901f\u5ea6\u63a5\u8fd1 5-12 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u5728 06:20 \u8dc3\u5347\u81f3 233.42 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 120.2%\u3002glm-5.3 \u8fd8\u5728 03:40 \u7684 184.0 \u8bcd\u5143/\u79d2\u4e0b\u964d\u81f3 05:40 \u7684 64.84 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0cvalid_point_count \u4e3a 96\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":3858559811,"translated_at":"2026-10-04T07:03:07.753995+00:00","translation_eval_count":229,"translation_model":"glm-5.3","translation_prompt_eval_count":368,"translation_total_duration_ns":2242283246,"translation_wall_seconds":2.404,"valid_point_count":96,"wall_seconds":4.03},{"coverage_pct":100.0,"eval_count":891,"generated_at":"2026-10-04T06:03:12.096837+00:00","generated_label":"Oct 04, 2026 \u00b7 06:03 UTC","id":1120,"model":"glm-5.3","period_end":"2026-10-04T06:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 02:03 UTC to Oct 04, 2026 \u00b7 06:03 UTC","period_start":"2026-10-04T02:03:01+00:00","prompt_eval_count":2249,"summary":"- gemma4:31b is the strongest model at 154.77 token/s average throughput, narrowly ahead of deepseek-v4.1-flash at 154.75 token/s; nemotron-3-ultra is the weakest at 36.89 token/s average.\n- glm-5.3-flash shows the most operationally significant volatility: after peaking at 182.04 token/s at 03:00, it fell below 12 token/s for seven consecutive observations (03:20 through 05:20), with a 127.9% coefficient of variation, before recovering to 130.12 token/s.\n- No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- gemma4:31b is the strongest model at 154.77 token/s average throughput, narrowly ahead of deepseek-v4.1-flash at 154.75 token/s; nemotron-3-ultra is the weakest at 36.89 token/s average.\n- glm-5.3-flash shows the most operationally significant volatility: after peaking at 182.04 token/s at 03:00, it fell below 12 token/s for seven consecutive observations (03:20 through 05:20), with a 127.9% coefficient of variation, before recovering to 130.12 token/s.\n- No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["gemma4:31b is the strongest model at 154.77 token/s average throughput, narrowly ahead of deepseek-v4.1-flash at 154.75 token/s; nemotron-3-ultra is the weakest at 36.89 token/s average.","glm-5.3-flash shows the most operationally significant volatility: after peaking at 182.04 token/s at 03:00, it fell below 12 token/s for seven consecutive observations (03:20 through 05:20), with a 127.9% coefficient of variation, before recovering to 130.12 token/s.","No missing-data limitation applies: all 8 models delivered 12 of 12 expected samples, yielding 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- gemma4:31b \u662f\u8868\u73b0\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 154.77 \u8bcd\u5143/\u79d2\uff0c\u4ee5\u5fae\u5f31\u4f18\u52bf\u9886\u5148\u4e8e 154.75 \u8bcd\u5143/\u79d2\u7684 deepseek-v4.1-flash\uff1bnemotron-3-ultra \u8868\u73b0\u6700\u5f31\uff0c\u5e73\u5747\u4e3a 36.89 \u8bcd\u5143/\u79d2\u3002\n- glm-5.3-flash \u7684\u6ce2\u52a8\u6027\u5bf9\u8fd0\u8425\u5f71\u54cd\u6700\u5927\uff1a\u5728 03:00 \u8fbe\u5230 182.04 \u8bcd\u5143/\u79d2\u7684\u5cf0\u503c\u540e\uff0c\u5b83\u5728\u8fde\u7eed\u4e03\u6b21\u89c2\u6d4b\uff0803:20 \u81f3 05:20\uff09\u4e2d\u8dcc\u7834 12 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 127.9%\uff0c\u968f\u540e\u6062\u590d\u81f3 130.12 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u7f3a\u5931\u6570\u636e\u7684\u9650\u5236\uff1a\u5168\u90e8 8 \u4e2a\u6a21\u578b\u5747\u4ea4\u4ed8\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5185\u5171\u4ea7\u751f 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":7742188583,"translated_at":"2026-10-04T06:03:12.096837+00:00","translation_eval_count":208,"translation_model":"glm-5.3","translation_prompt_eval_count":342,"translation_total_duration_ns":2806584729,"translation_wall_seconds":2.965,"valid_point_count":96,"wall_seconds":7.933},{"coverage_pct":100.0,"eval_count":1102,"generated_at":"2026-10-04T05:03:07.295735+00:00","generated_label":"Oct 04, 2026 \u00b7 05:03 UTC","id":1119,"model":"glm-5.3","period_end":"2026-10-04T05:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 01:03 UTC to Oct 04, 2026 \u00b7 05:03 UTC","period_start":"2026-10-04T01:03:01+00:00","prompt_eval_count":2251,"summary":"- deepseek-v4.1-flash is the strongest model at 165.19 token/s average throughput (peak 254.14 token/s), while nemotron-3-ultra is the weakest at 42.42 token/s average, peaking at only 76.96 token/s.\n- The most operationally significant volatility is glm-5.3-flash, which fell from a 191.44 token/s maximum to roughly 5-11 token/s in most observations after 02:20 UTC, with a 106.5% coefficient of variation and a -94.4% trend.\n- No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage for the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 165.19 token/s average throughput (peak 254.14 token/s), while nemotron-3-ultra is the weakest at 42.42 token/s average, peaking at only 76.96 token/s.\n- The most operationally significant volatility is glm-5.3-flash, which fell from a 191.44 token/s maximum to roughly 5-11 token/s in most observations after 02:20 UTC, with a 106.5% coefficient of variation and a -94.4% trend.\n- No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage for the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 165.19 token/s average throughput (peak 254.14 token/s), while nemotron-3-ultra is the weakest at 42.42 token/s average, peaking at only 76.96 token/s.","The most operationally significant volatility is glm-5.3-flash, which fell from a 191.44 token/s maximum to roughly 5-11 token/s in most observations after 02:20 UTC, with a 106.5% coefficient of variation and a -94.4% trend.","No missing-data limitation applies: all eight models report 12 of 12 expected samples, 96 valid points, and 100.0% coverage for the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 165.19 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c 254.14 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 42.42 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u4ec5\u4e3a 76.96 \u8bcd\u5143/\u79d2\u3002\n- \u8fd0\u8425\u5c42\u9762\u6700\u663e\u8457\u7684\u6ce2\u52a8\u51fa\u73b0\u5728 glm-5.3-flash\uff0c\u5176\u5728 UTC 02:20 \u4e4b\u540e\u7684\u5927\u591a\u6570\u89c2\u6d4b\u4e2d\u4ece 191.44 \u8bcd\u5143/\u79d2\u7684\u6700\u5927\u503c\u8dcc\u81f3\u7ea6 5-11 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 106.5%\uff0c\u8d8b\u52bf\u4e3a -94.4%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62a5\u544a\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\u300196 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u4ee5\u53ca\u56db\u5c0f\u65f6\u7a97\u53e3\u5185 100.0% \u7684\u8986\u76d6\u7387\u3002","total_duration_ns":4100068577,"translated_at":"2026-10-04T05:03:07.295735+00:00","translation_eval_count":190,"translation_model":"glm-5.3","translation_prompt_eval_count":336,"translation_total_duration_ns":927432427,"translation_wall_seconds":1.277,"valid_point_count":96,"wall_seconds":4.272},{"coverage_pct":100.0,"eval_count":1146,"generated_at":"2026-10-04T04:03:07.291740+00:00","generated_label":"Oct 04, 2026 \u00b7 04:03 UTC","id":1118,"model":"glm-5.3","period_end":"2026-10-04T04:03:01+00:00","period_label":"Oct 04, 2026 \u00b7 00:03 UTC to Oct 04, 2026 \u00b7 04:03 UTC","period_start":"2026-10-04T00:03:01+00:00","prompt_eval_count":2251,"summary":"- deepseek-v4.1-flash is the strongest model at 164.14 token/s average throughput, peaking at 254.14 token/s; minimax-m3 is the weakest at 49.33 token/s average, with a low of 10.31 token/s.\n- glm-5.3-flash shows the most severe volatility, with a 78.8% coefficient of variation and repeated collapses to roughly 7-12 token/s at 01:00, 02:20, 03:20, and 03:40 UTC; deepseek-v4.1-flash also swung from 254.14 token/s at 02:20 down to 21.39 token/s at 03:40 before recovering to 201.4 token/s.\n- No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- deepseek-v4.1-flash is the strongest model at 164.14 token/s average throughput, peaking at 254.14 token/s; minimax-m3 is the weakest at 49.33 token/s average, with a low of 10.31 token/s.\n- glm-5.3-flash shows the most severe volatility, with a 78.8% coefficient of variation and repeated collapses to roughly 7-12 token/s at 01:00, 02:20, 03:20, and 03:40 UTC; deepseek-v4.1-flash also swung from 254.14 token/s at 02:20 down to 21.39 token/s at 03:40 before recovering to 201.4 token/s.\n- No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["deepseek-v4.1-flash is the strongest model at 164.14 token/s average throughput, peaking at 254.14 token/s; minimax-m3 is the weakest at 49.33 token/s average, with a low of 10.31 token/s.","glm-5.3-flash shows the most severe volatility, with a 78.8% coefficient of variation and repeated collapses to roughly 7-12 token/s at 01:00, 02:20, 03:20, and 03:40 UTC; deepseek-v4.1-flash also swung from 254.14 token/s at 02:20 down to 21.39 token/s at 03:40 before recovering to 201.4 token/s.","No missing-data limitation applies: all eight models recorded 12 of 12 expected samples, with 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 164.14 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 254.14 \u8bcd\u5143/\u79d2\uff1bminimax-m3 \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 49.33 \u8bcd\u5143/\u79d2\uff0c\u6700\u4f4e\u4e3a 10.31 \u8bcd\u5143/\u79d2\u3002\n- glm-5.3-flash \u6ce2\u52a8\u6700\u4e3a\u5267\u70c8\uff0c\u53d8\u5f02\u7cfb\u6570\u8fbe 78.8%\uff0c\u5e76\u5728 UTC \u65f6\u95f4 01:00\u300102:20\u300103:20 \u548c 03:40 \u591a\u6b21\u9aa4\u964d\u81f3\u7ea6 7-12 \u8bcd\u5143/\u79d2\uff1bdeepseek-v4.1-flash \u4e5f\u4ece 02:20 \u7684 254.14 \u8bcd\u5143/\u79d2\u9aa4\u964d\u81f3 03:40 \u7684 21.39 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u56de\u5347\u81f3 201.4 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u8bb0\u5f55\u4e86 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u65f6\u95f4\u7a97\u53e3\u5185\u5171\u6709 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":4538294341,"translated_at":"2026-10-04T04:03:07.291740+00:00","translation_eval_count":231,"translation_model":"glm-5.3","translation_prompt_eval_count":373,"translation_total_duration_ns":1054265273,"translation_wall_seconds":1.207,"valid_point_count":96,"wall_seconds":4.712},{"coverage_pct":100.0,"eval_count":530,"generated_at":"2026-10-04T02:03:05.381382+00:00","generated_label":"Oct 04, 2026 \u00b7 02:03 UTC","id":1117,"model":"glm-5.3","period_end":"2026-10-04T02:03:01+00:00","period_label":"Oct 03, 2026 \u00b7 22:03 UTC to Oct 04, 2026 \u00b7 02:03 UTC","period_start":"2026-10-03T22:03:01+00:00","prompt_eval_count":2251,"summary":"- deepseek-v4.1-flash is the strongest model at 175.68 token/s average throughput, peaking at 233.3 token/s, while nemotron-3-ultra is the weakest at 40.98 token/s average, never exceeding 77.48 token/s.\n- The most operationally significant event is a synchronized throughput dip at 01:00 UTC, when glm-5.3-flash fell to 6.91 token/s, minimax-m3 to 10.31 token/s, and gemma4:31b to 36.03 token/s; minimax-m3 also shows the highest volatility with a 47.0% coefficient of variation.\n- No missing-data limitation applies: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.","summary_en":"- deepseek-v4.1-flash is the strongest model at 175.68 token/s average throughput, peaking at 233.3 token/s, while nemotron-3-ultra is the weakest at 40.98 token/s average, never exceeding 77.48 token/s.\n- The most operationally significant event is a synchronized throughput dip at 01:00 UTC, when glm-5.3-flash fell to 6.91 token/s, minimax-m3 to 10.31 token/s, and gemma4:31b to 36.03 token/s; minimax-m3 also shows the highest volatility with a 47.0% coefficient of variation.\n- No missing-data limitation applies: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented.","summary_items":["deepseek-v4.1-flash is the strongest model at 175.68 token/s average throughput, peaking at 233.3 token/s, while nemotron-3-ultra is the weakest at 40.98 token/s average, never exceeding 77.48 token/s.","The most operationally significant event is a synchronized throughput dip at 01:00 UTC, when glm-5.3-flash fell to 6.91 token/s, minimax-m3 to 10.31 token/s, and gemma4:31b to 36.03 token/s; minimax-m3 also shows the highest volatility with a 47.0% coefficient of variation.","No missing-data limitation applies: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0%, so the four-hour window is fully represented."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 175.68 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 233.3 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 40.98 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 77.48 \u8bcd\u5143/\u79d2\u3002\n- \u8fd0\u8425\u5c42\u9762\u6700\u91cd\u5927\u7684\u4e8b\u4ef6\u662f 01:00 UTC \u51fa\u73b0\u7684\u540c\u6b65\u541e\u5410\u91cf\u4e0b\u964d\uff0c\u5f53\u65f6 glm-5.3-flash \u964d\u81f3 6.91 \u8bcd\u5143/\u79d2\uff0cminimax-m3 \u964d\u81f3 10.31 \u8bcd\u5143/\u79d2\uff0cgemma4:31b \u964d\u81f3 36.03 \u8bcd\u5143/\u79d2\uff1bminimax-m3 \u8fd8\u8868\u73b0\u51fa\u6700\u9ad8\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 47.0%\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0cvalid_point_count \u4e3a 96\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u56e0\u6b64\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u5f97\u5230\u4e86\u5b8c\u6574\u5448\u73b0\u3002","total_duration_ns":2261112073,"translated_at":"2026-10-04T02:03:05.381382+00:00","translation_eval_count":219,"translation_model":"glm-5.3","translation_prompt_eval_count":360,"translation_total_duration_ns":817649212,"translation_wall_seconds":0.965,"valid_point_count":96,"wall_seconds":2.429},{"coverage_pct":100.0,"eval_count":925,"generated_at":"2026-10-04T01:03:07.134990+00:00","generated_label":"Oct 04, 2026 \u00b7 01:03 UTC","id":1116,"model":"glm-5.3","period_end":"2026-10-04T01:03:01+00:00","period_label":"Oct 03, 2026 \u00b7 21:03 UTC to Oct 04, 2026 \u00b7 01:03 UTC","period_start":"2026-10-03T21:03:01+00:00","prompt_eval_count":2250,"summary":"- deepseek-v4.1-flash is the strongest model at 173.88 token/s average throughput, peaking at 233.3 token/s; nemotron-3-ultra is weakest at 43.08 token/s average, never exceeding 77.48 token/s.\n- The most operationally significant volatility is the 01:00 collapse in glm-5.3-flash, which fell from 175.51 token/s at 00:20 to 6.91 token/s; deepseek-v4.1-flash also dipped to 33.76 token/s at 23:40 before recovering to 229.94 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage.","summary_en":"- deepseek-v4.1-flash is the strongest model at 173.88 token/s average throughput, peaking at 233.3 token/s; nemotron-3-ultra is weakest at 43.08 token/s average, never exceeding 77.48 token/s.\n- The most operationally significant volatility is the 01:00 collapse in glm-5.3-flash, which fell from 175.51 token/s at 00:20 to 6.91 token/s; deepseek-v4.1-flash also dipped to 33.76 token/s at 23:40 before recovering to 229.94 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage.","summary_items":["deepseek-v4.1-flash is the strongest model at 173.88 token/s average throughput, peaking at 233.3 token/s; nemotron-3-ultra is weakest at 43.08 token/s average, never exceeding 77.48 token/s.","The most operationally significant volatility is the 01:00 collapse in glm-5.3-flash, which fell from 175.51 token/s at 00:20 to 6.91 token/s; deepseek-v4.1-flash also dipped to 33.76 token/s at 23:40 before recovering to 229.94 token/s.","No missing-data limitation applies: all eight models have 12 of 12 samples, and the dataset reports 96 valid points with 100.0% coverage."],"summary_zh":"- deepseek-v4.1-flash \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 173.88 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 233.3 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 43.08 \u8bcd\u5143/\u79d2\uff0c\u4ece\u672a\u8d85\u8fc7 77.48 \u8bcd\u5143/\u79d2\u3002\n- \u8fd0\u8425\u5c42\u9762\u6700\u663e\u8457\u7684\u6ce2\u52a8\u662f glm-5.3-flash \u5728 01:00 \u7684\u9aa4\u964d\uff0c\u4ece 00:20 \u7684 175.51 \u8bcd\u5143/\u79d2\u8dcc\u81f3 6.91 \u8bcd\u5143/\u79d2\uff1bdeepseek-v4.1-flash \u4e5f\u5728 23:40 \u4e00\u5ea6\u964d\u81f3 33.76 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u6062\u590d\u81f3 229.94 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u6570\u636e\u96c6\u62a5\u544a 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":3936675566,"translated_at":"2026-10-04T01:03:07.134990+00:00","translation_eval_count":197,"translation_model":"glm-5.3","translation_prompt_eval_count":343,"translation_total_duration_ns":1051861081,"translation_wall_seconds":1.199,"valid_point_count":96,"wall_seconds":4.103},{"coverage_pct":100.0,"eval_count":306,"generated_at":"2026-10-04T00:03:08.835048+00:00","generated_label":"Oct 04, 2026 \u00b7 00:03 UTC","id":1115,"model":"glm-5.3","period_end":"2026-10-04T00:03:01+00:00","period_label":"Oct 03, 2026 \u00b7 20:03 UTC to Oct 04, 2026 \u00b7 00:03 UTC","period_start":"2026-10-03T20:03:01+00:00","prompt_eval_count":2249,"summary":"- gemma4:31b is the strongest model on average throughput at 143.0 token/s, ahead of deepseek-v4.1-flash at 164.21 token/s only on peaks (233.3 token/s maximum); glm-5.2 is the weakest at 43.18 token/s average, below nemotron-3-ultra at 42.0 token/s only marginally.\n- minimax-m3 shows the most operationally significant volatility, swinging between 11.52 and 105.05 token/s with 54.8% coefficient of variation and a 47.4% trend; deepseek-v4.1-flash also dropped to 33.76 token/s at 23:40.\n- No missing-data limitation exists: all 8 models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- gemma4:31b is the strongest model on average throughput at 143.0 token/s, ahead of deepseek-v4.1-flash at 164.21 token/s only on peaks (233.3 token/s maximum); glm-5.2 is the weakest at 43.18 token/s average, below nemotron-3-ultra at 42.0 token/s only marginally.\n- minimax-m3 shows the most operationally significant volatility, swinging between 11.52 and 105.05 token/s with 54.8% coefficient of variation and a 47.4% trend; deepseek-v4.1-flash also dropped to 33.76 token/s at 23:40.\n- No missing-data limitation exists: all 8 models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["gemma4:31b is the strongest model on average throughput at 143.0 token/s, ahead of deepseek-v4.1-flash at 164.21 token/s only on peaks (233.3 token/s maximum); glm-5.2 is the weakest at 43.18 token/s average, below nemotron-3-ultra at 42.0 token/s only marginally.","minimax-m3 shows the most operationally significant volatility, swinging between 11.52 and 105.05 token/s with 54.8% coefficient of variation and a 47.4% trend; deepseek-v4.1-flash also dropped to 33.76 token/s at 23:40.","No missing-data limitation exists: all 8 models report 12 of 12 samples, with 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- gemma4:31b \u5728\u5e73\u5747\u541e\u5410\u91cf\u65b9\u9762\u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u8fbe\u5230 143.0 \u8bcd\u5143/\u79d2\uff0c\u4ec5\u5728\u5cf0\u503c\u4e0a\u843d\u540e\u4e8e deepseek-v4.1-flash \u7684 164.21 \u8bcd\u5143/\u79d2\uff08\u6700\u9ad8 233.3 \u8bcd\u5143/\u79d2\uff09\uff1bglm-5.2 \u6700\u5f31\uff0c\u5e73\u5747 43.18 \u8bcd\u5143/\u79d2\uff0c\u4ec5\u7565\u5fae\u4f4e\u4e8e nemotron-3-ultra \u7684 42.0 \u8bcd\u5143/\u79d2\u3002\n- minimax-m3 \u8868\u73b0\u51fa\u6700\u5177\u8fd0\u8425\u610f\u4e49\u7684\u6ce2\u52a8\u6027\uff0c\u5728 11.52 \u81f3 105.05 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u6446\u52a8\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 54.8%\uff0c\u8d8b\u52bf\u4e3a 47.4%\uff1bdeepseek-v4.1-flash \u4e5f\u5728 23:40 \u964d\u81f3 33.76 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u7f3a\u5931\u6570\u636e\u7684\u9650\u5236\uff1a\u5168\u90e8 8 \u4e2a\u6a21\u578b\u5747\u62a5\u544a\u4e86 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u5171\u6709 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":3204802964,"translated_at":"2026-10-04T00:03:08.835048+00:00","translation_eval_count":461,"translation_model":"glm-5.3","translation_prompt_eval_count":362,"translation_total_duration_ns":3878383550,"translation_wall_seconds":4.2,"valid_point_count":96,"wall_seconds":3.368},{"coverage_pct":100.0,"eval_count":473,"generated_at":"2026-10-03T23:03:04.968472+00:00","generated_label":"Oct 03, 2026 \u00b7 23:03 UTC","id":1114,"model":"glm-5.3","period_end":"2026-10-03T23:03:01+00:00","period_label":"Oct 03, 2026 \u00b7 19:03 UTC to Oct 03, 2026 \u00b7 23:03 UTC","period_start":"2026-10-03T19:03:01+00:00","prompt_eval_count":2248,"summary":"- deepseek-v4.1-flash is the strongest model by average throughput at 163.79 token/s, while nemotron-3-ultra is the weakest at 42.45 token/s average.\n- glm-5.3 shows the most operationally significant volatility, with a coefficient of variation of 39.8% and swings between 31.52 and 182.87 token/s, alongside an 18.9% downward trend; by contrast, deepseek-v4.1-flash improved 26.3% over the window.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the full four-hour window is represented.","summary_en":"- deepseek-v4.1-flash is the strongest model by average throughput at 163.79 token/s, while nemotron-3-ultra is the weakest at 42.45 token/s average.\n- glm-5.3 shows the most operationally significant volatility, with a coefficient of variation of 39.8% and swings between 31.52 and 182.87 token/s, alongside an 18.9% downward trend; by contrast, deepseek-v4.1-flash improved 26.3% over the window.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the full four-hour window is represented.","summary_items":["deepseek-v4.1-flash is the strongest model by average throughput at 163.79 token/s, while nemotron-3-ultra is the weakest at 42.45 token/s average.","glm-5.3 shows the most operationally significant volatility, with a coefficient of variation of 39.8% and swings between 31.52 and 182.87 token/s, alongside an 18.9% downward trend; by contrast, deepseek-v4.1-flash improved 26.3% over the window.","No missing-data limitation applies: all eight models have 12 of 12 samples, valid_point_count is 96, and coverage is 100.0%, so the full four-hour window is represented."],"summary_zh":"- deepseek-v4.1-flash \u662f\u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u8fbe\u5230 163.79 \u8bcd\u5143/\u79d2\uff0c\u800c nemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 42.45 \u8bcd\u5143/\u79d2\u3002\n- glm-5.3 \u8868\u73b0\u51fa\u5bf9\u8fd0\u884c\u5f71\u54cd\u6700\u663e\u8457\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 39.8%\uff0c\u6ce2\u52a8\u8303\u56f4\u5728 31.52 \u81f3 182.87 \u8bcd\u5143/\u79d2\u4e4b\u95f4\uff0c\u5e76\u4f34\u968f 18.9% \u7684\u4e0b\u964d\u8d8b\u52bf\uff1b\u76f8\u6bd4\u4e4b\u4e0b\uff0cdeepseek-v4.1-flash \u5728\u8be5\u65f6\u95f4\u7a97\u53e3\u5185\u63d0\u5347\u4e86 26.3%\u3002\n- \u4e0d\u5b58\u5728\u7f3a\u5931\u6570\u636e\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0cvalid_point_count \u4e3a 96\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u56e0\u6b64\u5b8c\u6574\u8868\u793a\u4e86\u6574\u4e2a\u56db\u5c0f\u65f6\u7a97\u53e3\u3002","total_duration_ns":2117981523,"translated_at":"2026-10-03T23:03:04.968472+00:00","translation_eval_count":169,"translation_model":"glm-5.3","translation_prompt_eval_count":328,"translation_total_duration_ns":883314754,"translation_wall_seconds":1.029,"valid_point_count":96,"wall_seconds":2.281},{"coverage_pct":100.0,"eval_count":338,"generated_at":"2026-10-03T22:03:05.023824+00:00","generated_label":"Oct 03, 2026 \u00b7 22:03 UTC","id":1113,"model":"glm-5.3","period_end":"2026-10-03T22:03:01+00:00","period_label":"Oct 03, 2026 \u00b7 18:03 UTC to Oct 03, 2026 \u00b7 22:03 UTC","period_start":"2026-10-03T18:03:01+00:00","prompt_eval_count":2245,"summary":"- gemma4:31b is the strongest model with an average throughput of 140.28 token/s (peaking at 176.61 token/s), while nemotron-3-ultra is the weakest at 37.99 token/s average, never exceeding 66.59 token/s across the window.\n- deepseek-v4.1-flash shows the most operationally significant volatility, swinging from a low of 9.07 token/s at 18:40 to a high of 215.65 token/s at 21:40, with a 43.6% coefficient of variation and a 65.6% upward trend; glm-5.3 also oscillated between 31.52 and 182.87 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0% for the full four-hour period.","summary_en":"- gemma4:31b is the strongest model with an average throughput of 140.28 token/s (peaking at 176.61 token/s), while nemotron-3-ultra is the weakest at 37.99 token/s average, never exceeding 66.59 token/s across the window.\n- deepseek-v4.1-flash shows the most operationally significant volatility, swinging from a low of 9.07 token/s at 18:40 to a high of 215.65 token/s at 21:40, with a 43.6% coefficient of variation and a 65.6% upward trend; glm-5.3 also oscillated between 31.52 and 182.87 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0% for the full four-hour period.","summary_items":["gemma4:31b is the strongest model with an average throughput of 140.28 token/s (peaking at 176.61 token/s), while nemotron-3-ultra is the weakest at 37.99 token/s average, never exceeding 66.59 token/s across the window.","deepseek-v4.1-flash shows the most operationally significant volatility, swinging from a low of 9.07 token/s at 18:40 to a high of 215.65 token/s at 21:40, with a 43.6% coefficient of variation and a 65.6% upward trend; glm-5.3 also oscillated between 31.52 and 182.87 token/s.","No missing-data limitation applies: all eight models have 12 of 12 expected samples, valid_point_count is 96, and coverage is 100.0% for the full four-hour period."],"summary_zh":"- gemma4:31b \u662f\u6700\u5f3a\u7684\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 140.28 \u8bcd\u5143/\u79d2\uff08\u5cf0\u503c\u8fbe 176.61 \u8bcd\u5143/\u79d2\uff09\uff0c\u800c nemotron-3-ultra \u662f\u6700\u5f31\u7684\u6a21\u578b\uff0c\u5e73\u5747\u4e3a 37.99 \u8bcd\u5143/\u79d2\uff0c\u5728\u6574\u4e2a\u65f6\u95f4\u7a97\u53e3\u5185\u4ece\u672a\u8d85\u8fc7 66.59 \u8bcd\u5143/\u79d2\u3002\n- deepseek-v4.1-flash \u8868\u73b0\u51fa\u5bf9\u8fd0\u8425\u5f71\u54cd\u6700\u663e\u8457\u7684\u6ce2\u52a8\u6027\uff0c\u4ece 18:40 \u7684\u4f4e\u70b9 9.07 \u8bcd\u5143/\u79d2\u6446\u52a8\u5230 21:40 \u7684\u9ad8\u70b9 215.65 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 43.6%\uff0c\u4e0a\u5347\u8d8b\u52bf\u4e3a 65.6%\uff1bglm-5.3 \u4e5f\u5728 31.52 \u81f3 182.87 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u9707\u8361\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62e5\u6709 12 \u4e2a\u9884\u671f\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0cvalid_point_count \u4e3a 96\uff0c\u4e14\u5728\u6574\u4e2a\u56db\u5c0f\u65f6\u671f\u95f4\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":2007322071,"translated_at":"2026-10-03T22:03:05.023824+00:00","translation_eval_count":218,"translation_model":"glm-5.3","translation_prompt_eval_count":368,"translation_total_duration_ns":1084937180,"translation_wall_seconds":1.231,"valid_point_count":96,"wall_seconds":2.203},{"coverage_pct":100.0,"eval_count":611,"generated_at":"2026-10-03T21:03:06.395475+00:00","generated_label":"Oct 03, 2026 \u00b7 21:03 UTC","id":1112,"model":"glm-5.3","period_end":"2026-10-03T21:03:01+00:00","period_label":"Oct 03, 2026 \u00b7 17:03 UTC to Oct 03, 2026 \u00b7 21:03 UTC","period_start":"2026-10-03T17:03:01+00:00","prompt_eval_count":2249,"summary":"- Strongest average throughput was gemma4:31b at 142.86 token/s across 12 samples (range 96.61\u2013171.19 token/s); weakest was nemotron-3-ultra at 27.06 token/s across 12 samples (range 6.09\u201362.14 token/s).\n- deepseek-v4.1-flash showed the widest swings, spanning 9.07 to 231.01 token/s with a 54.4% coefficient of variation, while glm-5.3 rose 63.4% overall to a 174.8 token/s latest reading; deepseek-v4-pro also dipped to 37.21 token/s at 20:40 before recovering to 115.68 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, though the four-hour window limits longer-term conclusions.","summary_en":"- Strongest average throughput was gemma4:31b at 142.86 token/s across 12 samples (range 96.61\u2013171.19 token/s); weakest was nemotron-3-ultra at 27.06 token/s across 12 samples (range 6.09\u201362.14 token/s).\n- deepseek-v4.1-flash showed the widest swings, spanning 9.07 to 231.01 token/s with a 54.4% coefficient of variation, while glm-5.3 rose 63.4% overall to a 174.8 token/s latest reading; deepseek-v4-pro also dipped to 37.21 token/s at 20:40 before recovering to 115.68 token/s.\n- No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, though the four-hour window limits longer-term conclusions.","summary_items":["Strongest average throughput was gemma4:31b at 142.86 token/s across 12 samples (range 96.61\u2013171.19 token/s); weakest was nemotron-3-ultra at 27.06 token/s across 12 samples (range 6.09\u201362.14 token/s).","deepseek-v4.1-flash showed the widest swings, spanning 9.07 to 231.01 token/s with a 54.4% coefficient of variation, while glm-5.3 rose 63.4% overall to a 174.8 token/s latest reading; deepseek-v4-pro also dipped to 37.21 token/s at 20:40 before recovering to 115.68 token/s.","No missing-data limitation applies: all eight models have 12 of 12 samples, 96 valid points, and 100.0% coverage, though the four-hour window limits longer-term conclusions."],"summary_zh":"- \u5e73\u5747\u541e\u5410\u91cf\u6700\u5f3a\u7684\u662f gemma4:31b\uff0c12 \u4e2a\u6837\u672c\u5e73\u5747\u8fbe 142.86 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 96.61\u2013171.19 \u8bcd\u5143/\u79d2\uff09\uff1b\u6700\u5f31\u7684\u662f nemotron-3-ultra\uff0c12 \u4e2a\u6837\u672c\u5e73\u5747\u4e3a 27.06 \u8bcd\u5143/\u79d2\uff08\u8303\u56f4 6.09\u201362.14 \u8bcd\u5143/\u79d2\uff09\u3002\n- deepseek-v4.1-flash \u6ce2\u52a8\u5e45\u5ea6\u6700\u5927\uff0c\u8de8\u5ea6\u4ece 9.07 \u5230 231.01 \u8bcd\u5143/\u79d2\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 54.4%\uff0c\u800c glm-5.3 \u6574\u4f53\u4e0a\u5347 63.4%\uff0c\u6700\u65b0\u8bfb\u6570\u8fbe\u5230 174.8 \u8bcd\u5143/\u79d2\uff1bdeepseek-v4-pro \u4e5f\u5728 20:40 \u4e00\u5ea6\u8dcc\u81f3 37.21 \u8bcd\u5143/\u79d2\uff0c\u968f\u540e\u56de\u5347\u81f3 115.68 \u8bcd\u5143/\u79d2\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u6709 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5171 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\uff0c\u4f46\u56db\u5c0f\u65f6\u7684\u65f6\u95f4\u7a97\u53e3\u9650\u5236\u4e86\u5bf9\u66f4\u957f\u671f\u7ed3\u8bba\u7684\u63a8\u65ad\u3002","total_duration_ns":3227452842,"translated_at":"2026-10-03T21:03:06.395475+00:00","translation_eval_count":244,"translation_model":"glm-5.3","translation_prompt_eval_count":373,"translation_total_duration_ns":1123285931,"translation_wall_seconds":1.27,"valid_point_count":96,"wall_seconds":3.402},{"coverage_pct":100.0,"eval_count":859,"generated_at":"2026-10-03T20:03:11.842193+00:00","generated_label":"Oct 03, 2026 \u00b7 20:03 UTC","id":1111,"model":"glm-5.3","period_end":"2026-10-03T20:03:01+00:00","period_label":"Oct 03, 2026 \u00b7 16:03 UTC to Oct 03, 2026 \u00b7 20:03 UTC","period_start":"2026-10-03T16:03:01+00:00","prompt_eval_count":2249,"summary":"- gemma4:31b is the strongest model at 141.79 token/s average throughput, peaking at 163.56 token/s; nemotron-3-ultra is the weakest at 23.86 token/s average, ranging from 6.09 to 62.14 token/s.\n- deepseek-v4.1-flash shows the most operationally significant volatility, with a 53.2% coefficient of variation, swings from 231.01 token/s at 17:20 down to 9.07 token/s at 18:40, and a -33.2% trend; glm-5.3 also oscillates between 42.51 and 182.87 token/s.\n- No missing-data limitation applies: all eight models report 12 of 12 samples, giving 96 valid points and 100.0% coverage across the four-hour window.","summary_en":"- gemma4:31b is the strongest model at 141.79 token/s average throughput, peaking at 163.56 token/s; nemotron-3-ultra is the weakest at 23.86 token/s average, ranging from 6.09 to 62.14 token/s.\n- deepseek-v4.1-flash shows the most operationally significant volatility, with a 53.2% coefficient of variation, swings from 231.01 token/s at 17:20 down to 9.07 token/s at 18:40, and a -33.2% trend; glm-5.3 also oscillates between 42.51 and 182.87 token/s.\n- No missing-data limitation applies: all eight models report 12 of 12 samples, giving 96 valid points and 100.0% coverage across the four-hour window.","summary_items":["gemma4:31b is the strongest model at 141.79 token/s average throughput, peaking at 163.56 token/s; nemotron-3-ultra is the weakest at 23.86 token/s average, ranging from 6.09 to 62.14 token/s.","deepseek-v4.1-flash shows the most operationally significant volatility, with a 53.2% coefficient of variation, swings from 231.01 token/s at 17:20 down to 9.07 token/s at 18:40, and a -33.2% trend; glm-5.3 also oscillates between 42.51 and 182.87 token/s.","No missing-data limitation applies: all eight models report 12 of 12 samples, giving 96 valid points and 100.0% coverage across the four-hour window."],"summary_zh":"- gemma4:31b \u662f\u6700\u5f3a\u6a21\u578b\uff0c\u5e73\u5747\u541e\u5410\u91cf\u4e3a 141.79 \u8bcd\u5143/\u79d2\uff0c\u5cf0\u503c\u8fbe 163.56 \u8bcd\u5143/\u79d2\uff1bnemotron-3-ultra \u6700\u5f31\uff0c\u5e73\u5747\u4e3a 23.86 \u8bcd\u5143/\u79d2\uff0c\u6ce2\u52a8\u8303\u56f4\u5728 6.09 \u81f3 62.14 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u3002\n- deepseek-v4.1-flash \u8868\u73b0\u51fa\u5bf9\u8fd0\u8425\u5f71\u54cd\u6700\u5927\u7684\u6ce2\u52a8\u6027\uff0c\u53d8\u5f02\u7cfb\u6570\u4e3a 53.2%\uff0c\u4ece 17:20 \u7684 231.01 \u8bcd\u5143/\u79d2\u9aa4\u964d\u81f3 18:40 \u7684 9.07 \u8bcd\u5143/\u79d2\uff0c\u8d8b\u52bf\u4e3a -33.2%\uff1bglm-5.3 \u4e5f\u5728 42.51 \u81f3 182.87 \u8bcd\u5143/\u79d2\u4e4b\u95f4\u9707\u8361\u3002\n- \u4e0d\u5b58\u5728\u6570\u636e\u7f3a\u5931\u7684\u9650\u5236\uff1a\u5168\u90e8\u516b\u4e2a\u6a21\u578b\u5747\u62a5\u544a\u4e86 12 \u4e2a\u6837\u672c\u4e2d\u7684 12 \u4e2a\uff0c\u5728\u56db\u5c0f\u65f6\u7a97\u53e3\u5185\u5171\u63d0\u4f9b 96 \u4e2a\u6709\u6548\u6570\u636e\u70b9\uff0c\u8986\u76d6\u7387\u4e3a 100.0%\u3002","total_duration_ns":6945365336,"translated_at":"2026-10-03T20:03:11.842193+00:00","translation_eval_count":216,"translation_model":"glm-5.3","translation_prompt_eval_count":358,"translation_total_duration_ns":2999375682,"translation_wall_seconds":3.145,"valid_point_count":96,"wall_seconds":7.107}]}
