KubeBolt docs
GitHub

Insights Engine

24 built-in rules with a lifecycle — durable episodes with typed evidence, flap damping, an expiry watchdog, first-class mutes, and a per-rule policy layer.

The insights engine evaluates 24 built-in rules against every registered cluster. Each finding carries the affected resource, a human-readable message, and a specific suggestion with remediation steps or kubectl commands.

Since 2.1.0 a finding is no longer a light that is either on or off. It has a history, an owner and a policy: episodes record what happened, mutes record what you decided to tolerate, and the rule matrix records how loud each rule is allowed to be.

Rule definitions

8 Critical 14 Warning 2 Info
  • crash-loop critical

    Container in CrashLoopBackOff with more than 3 restarts

  • oom-killed critical

    Container terminated with OOMKilled (exit 137)

  • zero-replicas critical

    Deployment with 0 available replicas

  • node-not-ready critical

    Node condition Ready ≠ True

  • image-pull-backoff critical

    Pod in ImagePullBackOff state

  • progress-deadline-exceeded critical

    Rollout stalled past progressDeadlineSeconds

  • missing-config-dependency critical

    Container can't start: referenced ConfigMap/Secret not found

  • helm-release-failed critical

    Helm release left in a failed state after install/upgrade

  • cpu-throttle-risk warning

    CPU usage >80% of limit sustained

  • memory-pressure warning

    Memory usage >85% of limit

  • hpa-maxed-out warning

    HPA current replicas == max replicas

  • pvc-pending warning

    PVC in Pending state for >5 minutes

  • frequent-restarts warning

    Pod with >5 restarts in 24 hours (non crash-loop)

  • evicted-pods warning

    Pods evicted from node due to pressure

  • readiness-probe-failing warning

    Pod Running but not Ready for >2 minutes

  • liveness-probe-failing warning

    Liveness probe failing repeatedly before restarts pile up

  • service-no-endpoints warning

    Service has zero ready endpoints

  • policy-no-match warning

    NetworkPolicy podSelector matches no pods

  • pdb-no-match warning

    PodDisruptionBudget selector matches no pods

  • helm-release-hook-pending warning

    Helm release stuck pending >5 min (hook never completed)

  • cert-expiring warning

    cert-manager Certificate expired or within 14 days of expiry

  • argocd-out-of-sync warning

    ArgoCD Application OutOfSync or Degraded

  • resource-underrequest info

    Requests <40% of actual usage

  • policy-orphan info

    Namespace has running pods but no NetworkPolicy

Episodes

Every firing insight opens an episode — a durable row keyed on the fingerprint, with an append-only timeline of transitions: opened, flapped, severity moved, muted, unmuted, resolved, expired. Each transition records who caused it (system, rule:<id>, watchdog, user:<id>) and why.

An insight episode: the timeline, the recurrences with the same fingerprint, and the recommendation.
An insight episode: the timeline, the recurrences with the same fingerprint, and the recommendation. KubeBolt 2.1.1

Typed evidence. A rule records what it actually saw when it fired — the numbers and objects behind the verdict, not a re-read of the cluster minutes later. The episode keeps it, so the recommendation can be checked against the state that produced it even after the cluster has moved on. Kobi reads the same evidence through get_insight_episode.

Flap damping. A fingerprint that re-fires within 10 minutes of resolving is the same episode flapping — the flap counter goes up, no new episode is created, and no second notification is sent. Outside that window it is a genuinely new episode.

The expiry watchdog. A firing episode whose cluster stopped reporting is no longer verifiable — the agent left, the cluster was deleted. After 15 minutes without a signal (KUBEBOLT_INSIGHT_EXPIRE_TTL) the watchdog flips it to expired and it drops out of the Active list. expired is not resolved: nobody observed recovery, and the distinction is preserved in the timeline.

Resolutions are typed, so history reads as a narrative instead of a count: auto_recovered, remediated (an executed action cleared it, with the action id attached), manual, and rule_changed (a policy edit cleared it).

Active and History. The Insights page splits into two tabs. History has server-side filters (status, severity, rule, cluster, time window) and pagination, and each episode opens a detail page with its full timeline and its recurrence — how often this exact fingerprint has come back.

History answers for dead clusters. An episode keeps the cluster’s persisted display name, so a post-mortem on a cluster you already deleted still reads in names rather than UUIDs — which is precisely when the question arrives.

Retention rides the hourly maintenance pass with the insights horizon (KUBEBOLT_INSIGHTS_RETENTION_HORIZON, 7 days by default). Firing episodes are never pruned.

Mutes

A mute silences one rule on one resource in one cluster. It is a display layer, not a suppression: the engine keeps evaluating and episodes keep recording — only your attention is spared.

Re-muting the same (cluster, rule, resource) key updates the terms instead of stacking duplicates.

The shift report

Home’s greeting tells you what happened while you were away. The window hangs from your own presence anchor — the last time the Home page actually rendered for you — not from your last login, which token refresh would freeze. Your first shift defaults to a 24-hour window; an anchor older than 30 days is clamped, and the report says so.

It reports events, not standing state: what opened, what auto-recovered, what was remediated, what expired, how many of the newly opened episodes are still burning, and how many were critical. A condition that merely spans your absence belongs to the Active list, not here — so the report goes quiet when nothing actually happened.

Bursts are clustered deterministically — the same window in produces the same grouping out — into node_rotation, node_pressure, mass_rollout, or an honest unknown_burst when nothing recognizable fits. The longest episode of the window is named and linked, and the mute counters (“silenced while you were away”, “silences standing now”) are part of the story.

The rule matrix

Rules split into two classes, and the class decides which knob you get:

ClassCountWhat you can change
Malfunction15Only the numeric bar, and only on rules that have one. Severity is locked and the rule can never be switched off — broken is broken anywhere. The only full silence for one resource is a mute.
Expectation9Severity is the dial, off included. These are violated on purpose in a development cluster.

Six rules carry a numeric threshold: crash-loop (restarts), frequent-restarts (restarts), resource-underrequest (request/usage ratio), cpu-throttle-risk and memory-pressure (usage/limit ratio), and cert-expiring (days before expiry). Moving the bar keeps the rule meaning what its name says; the other eighteen are boolean and have no dial.

Administration → Insights → Rules lists all 24 with their class, a one-sentence description, their knobs, and the honesty column Ignored 30d — episodes of that rule closed in the last 30 days without an acknowledgement, an action, or a mute. A rule you have turned off is still evaluated and counted: the Insights page reports how many findings your policy is hiding rather than pretending they did not happen.

The global override layer — one setting per rule, across every cluster — ships in the open-source distribution. Layering a second set of overrides per environment category (production, staging, testing, development) needs the cluster classification and is a KubeBolt Cloud feature; there, the effective value resolves knob by knob as environment override → global override → shipped default.

Endpoints

MethodEndpointDescription
GET/insightsActive insights for the current cluster, with hiddenByProfile and profile
GET/insights/summaryPer-cluster state for the fleet, served without a live connector
GET/insights/episodesEpisode history — status, severity, rule, cluster, since, until, limit, page
GET/insights/episodes/{id}One episode with its transitions and recurrence
GET/insights/operational-episodesThe deterministic bursts over a window
GET/insights/shift-report”While you were away”, scoped to your presence anchor
POST/account/dashboard-seenThe presence beacon Home fires after rendering
GET · POST/insights/mutesList (any role) and create (Editor, audited)
DELETE/insights/mutes/{id}Lift a mute (Editor, audited)
GET · PUT · DELETE/admin/insight-policies[/{rule}]The rule policy layer (Admin, audited)

None of these require a connected cluster — history has to answer for clusters that are gone.

Configuration

VariableDefaultPurpose
KUBEBOLT_INSIGHT_EXPIRE_TTL15mA firing episode with no signal beyond this flips to expired
KUBEBOLT_INSIGHTS_RETENTION_HORIZON168h (7 days)How far back resolved and expired episodes are kept

Insight notifications are configured separately — see Notifications.