mirror of
https://github.com/modular/skills.git
synced 2026-09-19 03:30:35 +08:00
a901eeb211
The serving benchmark console now prints GPU memory and utilization through the same percentile table as latency and throughput, instead of three lines per GPU. The section title states that the rows are aggregated across N devices, and a note points at the result JSON for per-GPU series. Those lists stay in result_groups.gpu_stats, so machine consumers are unchanged. MODULAR_ORIG_COMMIT_REV_ID: e0141c249236161d66a1b4dc17df7943c23f6217