Skip to content
Home

Proxmox Observability

Your Proxmox datacenter, as a navigable model.

OpenTelemetry-native metrics for your Proxmox VE estate, arranged as a datacenter → node → guest hierarchy with contextual interlinking, PSI-first pressure visibility, and host-vs-guest correlation. Full cardinality across qemu and LXC, no per-node tax, no proprietary agent — scrape the Proxmox VE Exporter straight into your Collector.

Seven Dashboards - One Datacenter Model

A navigable model, not a wall of graphs.

Click any card to enlarge · esc to close

Cardinal Proxmox Datacenter dashboard: Overview tiles (Nodes, Guests, Running, Storage Used, Storage Capacity, Guests Under Pressure) above a Needs Attention section listing guests by peak PSI % alongside stopped guests.

01 · Datacenter overview

Every node, every guest, one screen.

Inventory tiles for nodes, guests, and storage above a Needs Attention list of what's actually wrong — stopped guests, PSI-bound VMs, pools past threshold — each with the exact metric that surfaced it. No opaque health score.

Cardinal Proxmox Datacenter dashboard: Overview tiles (Nodes, Guests, Running, Storage Used, Storage Capacity, Guests Under Pressure) above a Needs Attention section listing guests by peak PSI % alongside stopped guests.
Contextual interlink menu opened on a stacked-bar chart of Node Fleet 5m load average, offering Investigate options — Open node pve40, Resource pressure on pve40, Network for pve40 — plus Related shortcuts to Guests on pve40 and Storage on pve40.

02 · Drill-down in three clicks

Every red thing is a link.

Click a hot node, land on its deep dive. Click a contended guest, land on it with the host and neighbor VMs already loaded. Time range and entity context carry across every jump — no PromQL, no refilter.

Contextual interlink menu opened on a stacked-bar chart of Node Fleet 5m load average, offering Investigate options — Open node pve40, Resource pressure on pve40, Network for pve40 — plus Related shortcuts to Guests on pve40 and Storage on pve40.
Cardinal Proxmox Node deep dive: Overview tiles for CPU, Memory, Load (5m), Guests, Network RX, Network TX, Uptime, Cores, and RAM Total, above a Resource Saturation section with Node CPU % and Node Memory % time series.

03 · Node deep dive

Which workloads are pressuring this host?

Per-node saturation, capacity-vs-allocation, and a guest contention table sorted by PSI — so the VMs actually driving the host rise to the top, right next to the host metrics themselves.

Cardinal Proxmox Node deep dive: Overview tiles for CPU, Memory, Load (5m), Guests, Network RX, Network TX, Uptime, Cores, and RAM Total, above a Resource Saturation section with Node CPU % and Node Memory % time series.
Cardinal Proxmox Guest — VM/LXC deep dive: Overview tiles for CPU, Memory, CPU Pressure Stall, Memory Pressure Stall, IO Pressure Stall, Disk IO, Network, and Uptime, above a Resource Pressure section with CPU Utilization and CPU Pressure Stall (some / full) time series.

04 · Guest + host context

Is it the VM, or the host it's on?

Every guest metric shown side-by-side with the same metric on its host and on the neighbour guests sharing the node — the virtualization question answered without switching dashboards.

Cardinal Proxmox Guest — VM/LXC deep dive: Overview tiles for CPU, Memory, CPU Pressure Stall, Memory Pressure Stall, IO Pressure Stall, Disk IO, Network, and Uptime, above a Resource Pressure section with CPU Utilization and CPU Pressure Stall (some / full) time series.
Cardinal Proxmox Resource Pressure dashboard: stacked-bar rankings of CPU-Constrained Guests by CPU Pressure Stall (some) alongside Memory-Constrained Guests by Memory Pressure Stall (some) plus swap flag, filtered across every node.

05 · PSI-first pressure

Pressure Stall, not just utilisation.

CPU, memory, IO, and node load ranked by PSI across the whole fleet. 80% CPU is fine until workloads start waiting — PSI is where that waiting actually shows up.

Cardinal Proxmox Resource Pressure dashboard: stacked-bar rankings of CPU-Constrained Guests by CPU Pressure Stall (some) alongside Memory-Constrained Guests by Memory Pressure Stall (some) plus swap flag, filtered across every node.
Cardinal Proxmox Storage dashboard: Utilization tiles (utilisation %, used GiB, free GiB, total capacity) above a Utilization gauge and a Utilization Over Time chart showing per-pool used percentage.

06 · Storage + growth

Every pool, plus days-to-full.

Per-pool used % and free GiB tiles for every pool on every node, utilisation over time, growth per week, and a projection of days-to-90% at current rate — interlinked to the guests driving the IO.

Cardinal Proxmox Storage dashboard: Utilization tiles (utilisation %, used GiB, free GiB, total capacity) above a Utilization gauge and a Utilization Over Time chart showing per-pool used percentage.
Cardinal Proxmox Network dashboard: Overview tiles for Guests in Scope, Total Guest RX (avg), Total Guest TX (avg), Total Node RX (avg), and Total Node TX (avg), above a Network section with Node NIC Receive and Transmit throughput time series.

07 · Network top talkers

Guest RX/TX next to host RX/TX.

Per-node NIC throughput and per-guest RX/TX ranked. Enter from a guest or a node — the top talkers on that surface load with the same time window, so a noisy neighbour is one click away.

Cardinal Proxmox Network dashboard: Overview tiles for Guests in Scope, Total Guest RX (avg), Total Guest TX (avg), Total Node RX (avg), and Total Node TX (avg), above a Network section with Node NIC Receive and Transmit throughput time series.
Cardinal Proxmox Availability & SLO dashboard: Fleet Uptime SLO, Guests Running, and Guests Stopped tiles, a Guests by Type donut (qemu vs lxc), Running Guest Count and Fleet Availability % time series, and a Per-Guest Window SLO grid.

08 · Availability & SLO

Every VM on its own uptime SLO.

Fleet availability, running-vs-stopped counts, worst-guest SLO surfaced first, and per-guest uptime targets — the specialist view for SLO owners, kept out of the default operational dashboard.

Cardinal Proxmox Availability & SLO dashboard: Fleet Uptime SLO, Guests Running, and Guests Stopped tiles, a Guests by Type donut (qemu vs lxc), Running Guest Count and Fleet Availability % time series, and a Per-Guest Window SLO grid.

Point your Collector at the Proxmox VE Exporter. Cardinal turns it into a navigable datacenter.

Scrape the Proxmox VE Exporter and your PVE node exporters into an OpenTelemetry Collector that writes to your own S3, GCS, or Azure Blob bucket — no vendor storage in the path. Seven interlinked dashboards, per-guest SLOs, qemu- and LXC-level depth, and cluster rollups are on the moment the first object lands.

  1. Step 1

    Create an account or log in to get started.

    Spin up a Cardinal workspace in under a minute. Free to start, no credit card.

    Create account / Log in
  2. Step 2

    Start sending your data — follow the guide.

    Point your OpenTelemetry Collector at the Proxmox VE Exporter and ship metrics to your own object store. The docs walk you through the config.

    Follow the guide
Compare vendors →