Insights Engine
24 built-in rules with a lifecycle — durable episodes with typed evidence, flap damping, an expiry watchdog, first-class mutes, and a per-rule policy layer.
The insights engine evaluates 24 built-in rules against every registered
cluster. Each finding carries the affected resource, a human-readable message,
and a specific suggestion with remediation steps or kubectl commands.
Since 2.1.0 a finding is no longer a light that is either on or off. It has a history, an owner and a policy: episodes record what happened, mutes record what you decided to tolerate, and the rule matrix records how loud each rule is allowed to be.
Rule definitions
- crash-loop critical
Container in CrashLoopBackOff with more than 3 restarts
- oom-killed critical
Container terminated with OOMKilled (exit 137)
- zero-replicas critical
Deployment with 0 available replicas
- node-not-ready critical
Node condition Ready ≠ True
- image-pull-backoff critical
Pod in ImagePullBackOff state
- progress-deadline-exceeded critical
Rollout stalled past progressDeadlineSeconds
- missing-config-dependency critical
Container can't start: referenced ConfigMap/Secret not found
- helm-release-failed critical
Helm release left in a failed state after install/upgrade
- cpu-throttle-risk warning
CPU usage >80% of limit sustained
- memory-pressure warning
Memory usage >85% of limit
- hpa-maxed-out warning
HPA current replicas == max replicas
- pvc-pending warning
PVC in Pending state for >5 minutes
- frequent-restarts warning
Pod with >5 restarts in 24 hours (non crash-loop)
- evicted-pods warning
Pods evicted from node due to pressure
- readiness-probe-failing warning
Pod Running but not Ready for >2 minutes
- liveness-probe-failing warning
Liveness probe failing repeatedly before restarts pile up
- service-no-endpoints warning
Service has zero ready endpoints
- policy-no-match warning
NetworkPolicy podSelector matches no pods
- pdb-no-match warning
PodDisruptionBudget selector matches no pods
- helm-release-hook-pending warning
Helm release stuck pending >5 min (hook never completed)
- cert-expiring warning
cert-manager Certificate expired or within 14 days of expiry
- argocd-out-of-sync warning
ArgoCD Application OutOfSync or Degraded
- resource-underrequest info
Requests <40% of actual usage
- policy-orphan info
Namespace has running pods but no NetworkPolicy
Episodes
Every firing insight opens an episode — a durable row keyed on the
fingerprint, with an append-only timeline of transitions: opened, flapped,
severity moved, muted, unmuted, resolved, expired. Each transition records who
caused it (system, rule:<id>, watchdog, user:<id>) and why.
Typed evidence. A rule records what it actually saw when it fired — the
numbers and objects behind the verdict, not a re-read of the cluster minutes
later. The episode keeps it, so the recommendation can be checked against the
state that produced it even after the cluster has moved on. Kobi reads the same
evidence through get_insight_episode.
Flap damping. A fingerprint that re-fires within 10 minutes of resolving is the same episode flapping — the flap counter goes up, no new episode is created, and no second notification is sent. Outside that window it is a genuinely new episode.
The expiry watchdog. A firing episode whose cluster stopped reporting is no
longer verifiable — the agent left, the cluster was deleted. After
15 minutes without a signal (KUBEBOLT_INSIGHT_EXPIRE_TTL) the watchdog
flips it to expired and it drops out of the Active list. expired is not
resolved: nobody observed recovery, and the distinction is preserved in the
timeline.
Resolutions are typed, so history reads as a narrative instead of a count:
auto_recovered, remediated (an executed action cleared it, with the action
id attached), manual, and rule_changed (a policy edit cleared it).
Active and History. The Insights page splits into two tabs. History has
server-side filters (status, severity, rule, cluster, time window) and
pagination, and each episode opens a detail page with its full timeline and its
recurrence — how often this exact fingerprint has come back.
History answers for dead clusters. An episode keeps the cluster’s persisted display name, so a post-mortem on a cluster you already deleted still reads in names rather than UUIDs — which is precisely when the question arrives.
Retention rides the hourly maintenance pass with the insights horizon
(KUBEBOLT_INSIGHTS_RETENTION_HORIZON, 7 days by default). Firing episodes are
never pruned.
Mutes
A mute silences one rule on one resource in one cluster. It is a display layer, not a suppression: the engine keeps evaluating and episodes keep recording — only your attention is spared.
- Always bounded: an expiry date, or until resolved. A permanent mute is possible but demands a written reason, so whoever finds it in six months knows why it exists.
- Mutes subtract everywhere insights are counted, the Overview KPI included.
The Insights page is the only surface that can show them back, behind
includeMuted=true. - Critical severity pierces a mute. An escalation gets through, which is why the engine never consults the mute list — piercing needs the underlying signal intact.
- Every mute and unmute lands in the episode timeline and the audit trail with the actor’s name and the reason.
- Reading the list needs no particular role; creating or lifting a mute needs Editor.
- Administration → Insights → Silenced lists them with cluster names, not UUIDs.
Re-muting the same (cluster, rule, resource) key updates the terms instead of
stacking duplicates.
The shift report
Home’s greeting tells you what happened while you were away. The window hangs from your own presence anchor — the last time the Home page actually rendered for you — not from your last login, which token refresh would freeze. Your first shift defaults to a 24-hour window; an anchor older than 30 days is clamped, and the report says so.
It reports events, not standing state: what opened, what auto-recovered, what was remediated, what expired, how many of the newly opened episodes are still burning, and how many were critical. A condition that merely spans your absence belongs to the Active list, not here — so the report goes quiet when nothing actually happened.
Bursts are clustered deterministically — the same window in produces the
same grouping out — into node_rotation, node_pressure, mass_rollout, or an
honest unknown_burst when nothing recognizable fits. The longest episode of
the window is named and linked, and the mute counters (“silenced while you were
away”, “silences standing now”) are part of the story.
The rule matrix
Rules split into two classes, and the class decides which knob you get:
| Class | Count | What you can change |
|---|---|---|
| Malfunction | 15 | Only the numeric bar, and only on rules that have one. Severity is locked and the rule can never be switched off — broken is broken anywhere. The only full silence for one resource is a mute. |
| Expectation | 9 | Severity is the dial, off included. These are violated on purpose in a development cluster. |
Six rules carry a numeric threshold: crash-loop (restarts), frequent-restarts
(restarts), resource-underrequest (request/usage ratio), cpu-throttle-risk
and memory-pressure (usage/limit ratio), and cert-expiring (days before
expiry). Moving the bar keeps the rule meaning what its name says; the other
eighteen are boolean and have no dial.
Administration → Insights → Rules lists all 24 with their class, a one-sentence description, their knobs, and the honesty column Ignored 30d — episodes of that rule closed in the last 30 days without an acknowledgement, an action, or a mute. A rule you have turned off is still evaluated and counted: the Insights page reports how many findings your policy is hiding rather than pretending they did not happen.
The global override layer — one setting per rule, across every cluster — ships in the open-source distribution. Layering a second set of overrides per environment category (production, staging, testing, development) needs the cluster classification and is a KubeBolt Cloud feature; there, the effective value resolves knob by knob as environment override → global override → shipped default.
Endpoints
| Method | Endpoint | Description |
|---|---|---|
| GET | /insights | Active insights for the current cluster, with hiddenByProfile and profile |
| GET | /insights/summary | Per-cluster state for the fleet, served without a live connector |
| GET | /insights/episodes | Episode history — status, severity, rule, cluster, since, until, limit, page |
| GET | /insights/episodes/{id} | One episode with its transitions and recurrence |
| GET | /insights/operational-episodes | The deterministic bursts over a window |
| GET | /insights/shift-report | ”While you were away”, scoped to your presence anchor |
| POST | /account/dashboard-seen | The presence beacon Home fires after rendering |
| GET · POST | /insights/mutes | List (any role) and create (Editor, audited) |
| DELETE | /insights/mutes/{id} | Lift a mute (Editor, audited) |
| GET · PUT · DELETE | /admin/insight-policies[/{rule}] | The rule policy layer (Admin, audited) |
None of these require a connected cluster — history has to answer for clusters that are gone.
Configuration
| Variable | Default | Purpose |
|---|---|---|
KUBEBOLT_INSIGHT_EXPIRE_TTL | 15m | A firing episode with no signal beyond this flips to expired |
KUBEBOLT_INSIGHTS_RETENTION_HORIZON | 168h (7 days) | How far back resolved and expired episodes are kept |
Insight notifications are configured separately — see Notifications.