@rodrimora
Post
Post 1 of 2
In case anyone wants a Grafana dashboard to monitor their DGX Cluster, I uploaded mine here https://github.com/RodriMora/dgx-spark-grafana-dashboard
Image from X post
Post 2 of 2
It has extra stuff like temps monitoring and thermal throttling. Point your agent at your Grafana/Prometheus instance and to your DGX cluster and it should be able to have it ready in no time.
Explanation
What it says @rodrimora published his Grafana dashboard for monitoring an NVIDIA DGX cluster, with the configuration on GitHub. Besides basic utilization/throughput, it includes temperature and thermal-throttling monitoring. His suggested setup path is to give an agent access to the Grafana/Prometheus instance plus the DGX cluster and have it configure the dashboard.
Context This is operational tooling rather than a benchmark claim: a prebuilt observability layer for DGX systems using Prometheus metrics visualized in Grafana. The screenshot appears specifically tuned around inference economics and performance, not merely hardware health.
Why it matters Useful if running DGX Spark/cluster inference and you want one screen showing whether performance problems come from workload behavior, caching, contention, memory pressure, or thermals. The particularly useful combination is latency + throughput + speculative-decoding + resource metrics, which makes tuning changes measurable rather than anecdotal.
Images The dashboard shows ~2,448 total tokens/s and 44.9 output tokens/s at the captured moment, 97.3% prefix-cache effectiveness, ~57.4% speculative-decode acceptance, ~92% RAM utilization, TTFT percentile traces, request queue state, disk/CPU utilization, and acceptance by speculative position (85.7% → 43.2%). It also estimates comparative cloud cost versus DGX energy cost.