Grafana
Dashboards and alerting on top of Prometheus — how teams actually see production health.
Maturity level: L2 — Can Build
Six perspectives on Grafana
Roadmap
Learn immediately after Prometheus — they are a pair in production.
Architecture
Visualization and alerting layer on metrics, logs, and traces.
Company
Prometheus/Grafana stack appears frequently in SRE, MLOps, and platform roles.
Projects
Every deployed service gets a dashboard before it goes live.
Interview
Metrics selection, SLOs, and on-call practices.
Career
Required for anyone owning production reliability.
What & Why
What: Observability platform for metrics visualization, dashboards, and alerting.
Why: Prometheus collects metrics; Grafana is what on-call engineers stare at. Job posts list both together.
Build this
Grafana dashboard for an ML API: latency, error rate, GPU memory, and model version.
Production reality
- ! Alert fatigue
- ! Missing labels
- ! Dashboard sprawl
- ! Stale panels
- ! False positives
Interview preparation
- What metrics matter for ML inference?
- How do you design actionable alerts?
- SLO-based alerting
Connected skills
Explore Grafana in the interactive universe or train with live cohorts.