KubeBolt docs
GitHub

Dashboard: Overview · Capacity · Reliability · Cost

The per-cluster dashboard follows a scan-then-dive rhythm — a summary layer that answers "is anything on fire?", with deep instrumentation below.

Each cluster’s dashboard is organized into sub-tabs built around a scan → dive rhythm: a summary layer up top answers the urgent question at a glance, and the detailed panels below let you drill into anything. Overview and Capacity are always there; Reliability appears when the cluster ships Hubble flow data, and Cost (beta) when it ships OpenCost data. A tab with nothing to show is hidden rather than rendered empty.

Sub-tabQuestion it answersNeeds
OverviewIs everything fine right now?Nothing beyond a connection; Metrics Server for live usage
CapacityHow is the cluster consuming, and is it sized right?Metrics history (the agent) for the trend charts
ReliabilityWhat is the cluster actually serving?Agent with Hubble enabled (Cilium)
Cost (beta)What does it cost, and what is idle?OpenCost data — see Cost

The card row every tab opens with

Overview, Capacity, Reliability and Cost each open with a row of four cards built the same way: a large figure, one sentence saying what it means right now, a small chart chosen for that metric, and a mono caption with the breakdown and links to the lists behind it.

Overview — is everything fine right now?

Overview has no range selector: everything on it is current state, and the trends live in Capacity. Its four cards:

CardShowsLights up
Cluster healthThe score out of 100 on a half gauge. It starts from the component checks and takes off 5 points per open critical insight (at most 25) and 2 per warning (at most 10); the card lists what took points off, and hovering shows every checkRed while a critical insight is open
Nodes readyReady nodes over the total, and the three busiest (by the higher of CPU and memory) as double rings — CPU inside, memory outside — with “+N more”. Without node metrics, a Ready / Not ready listRed when a node is not ready
PodsRunning pods over the active ones, and the pod count over the last 24 hours as a line (needs the agent; without it, a Running / Degraded list). The caption splits completed · degraded · not running, each linking to its filtered listRed when pods are not running, amber when they are degraded (Pending, CrashLoopBackOff, partially ready)
InsightsOpen insights, the three rules with the most of them (worst severity first), and the critical · warning · info splitRed with a critical open, amber with a warning

Completed Job pods are left out of the pods denominator, so a batch of finished Jobs never reads as missing pods.

Below the cards:

The Overview tab: efficiency, workload health and what needs attention, on one screen.
The Overview tab: efficiency, workload health and what needs attention, on one screen. KubeBolt 2.1.1
Capacity: requests against real usage, per node and per workload.
Capacity: requests against real usage, per node and per workload. KubeBolt 2.1.1

Reliability — golden signals, honest colors

Reliability: the golden signals, with colors that only turn when something is actually wrong.
Reliability: the golden signals, with colors that only turn when something is actually wrong. KubeBolt 2.1.1

The Reliability tab is powered by Hubble flow data, so it appears only on agent-connected clusters with Hubble enabled (hubble.enabled=true on the agent chart). See Connecting clusters.

Shared behavior

The selected time range is shared between Capacity, Reliability and Cost and persists while you drill into resources, but resets to the 15-minute default on a fresh browser session — the daily scan never accidentally fires a 30-day query.

Cost (beta)

Spend, idle cost, cost per pod and money-priced rightsizing, from OpenCost data the agent ships. It is covered on its own page: Cost.