How it works · the incident loop
It is 3am and something breaks. Here is what happens next.
The same incident the home page resolves in 87 seconds, step by step: what detects it, what the model reads, what it may touch and what is written down by morning.
01 — Detect · L1
Nobody pages you. It already noticed.
Twenty-four deterministic rules watch the cluster's state: crash loops, OOM kills, images that will not pull, certificates about to expire, nodes going away. Zero configuration, and not one AI credit. And they do not hand you a bare finding: each one arrives with its recommendation, so you can fix it yourself even when that incident is not in the catalog Autopilot handles on its own.
The six layers an incident travels through
Detector
rules over events
100% deterministicRouter
classifies and prioritizes
AIInvestigator
correlates logs, metrics and deploys
AIPlanner
builds the plan
AIExecutor
executes behind a deterministic guardrail
AI + deterministicPostmortem
timeline and evidence
AIEach chapter below is one of them. Only the first is 100% deterministic; the last one that touches production works behind a guardrail that is deterministic too.
L1 Insights Engine (L1): 24 zero-config rules detect the incident before any model gets involved. See the 24 rules →
- The rules run on current state, not on history, so they work from the first minute and do not depend on how much retention your plan has.
- Detecting costs no credits. They are spent only when something needs investigating, and the next layer makes that call.
- Every rule is a Skill: it carries its own confidence threshold and says which catalog actions are valid for that incident.
02 — Diagnose · L2 to L4
The model proposes. It never executes.
When that signal becomes a question, Kobi reasons and proposes the answer, but doesn't execute it: you authorize, and it's all audited.
Backend
runs the tools
results
LLM
AI provider
proposal
Card
human approval
only after approval
Your cluster
agent RBAC
The tools run in the backend. The model only sees results.
- BYOK (your own key or endpoint: Azure OpenAI, vLLM, Ollama) applies to OSS and EE Self-Hosted: AI traffic goes from your install to your provider, never through KubeBolt. On SaaS the AI is managed.
- You authorize each action and it runs with the agent's RBAC; KubeBolt never impersonates your credentials or your Kubernetes RBAC. Who gets read-only or edit is set by your org admin, and every action is audited: who authorized it, when, and which resources it touched.
- A tampered log can't execute anything. Logs are text a third party may have written; that's why the model proposes instead of executing.
The rule says what broke; why it broke needs context: events, logs, metrics and recent changes. It is the only thing that crosses the channel.
The agent opens the connection. The cluster exposes no port.
- Secret values are redacted before they leave the API layer — not a prompt policy, a server-side check.
- Metadata, not content. Each flow carries source, destination, verdict, status code and latency — never the request body.
- Logs aren't shipped in bulk. They're read only when you ask for a diagnosis that needs them.
- Self-hosted, nothing reaches KubeBolt: no product telemetry, no usage ping.
03 — Remediate · L5
Kobi acts alone. You approve anything destructive.
With nobody to ask, the signal fires Kobi: it detects, investigates, remediates and documents. The model proposes; a deterministic guardrail checks the action against the catalog before anything touches production.
- live
Anything destructive is approved by a person
A rollback always requires approval, at any autonomy level.
- live
If verification fails, Autopilot reverts
The change goes back to its previous state, automatically.
- live
Every action is logged
With its result and who approved it.
Three autonomy levels: suggest only (the default), approve and execute, autonomous. And off, if you want. With a confidence gate and per-cluster blocked namespaces.
If a step isn't on the list, the plan is rejected and escalates to a person.
- The plan is validated against code, not against a prompt. A new action needs a release, not a different instruction.
- An unattended window (days, hours, timezone) and a platform-level autonomy ceiling: outside the window, the incident is recorded, not acted on.
04 — Verify and revert
If the fix does not hold, it undoes it and calls you.
Applying is not resolving. Autopilot captures the previous values before touching anything, checks the result and, if it does not pass, restores and escalates to a person. It ships today; it is not on the roadmap.
Captures the previous state
The values it is about to change are captured before anything is applied, so the revert restores exactly what was there.
Checks the result
It waits for pods to go Ready and to stop restarting. An action applying without error is not enough.
Reverts if it fails
In autonomous mode, if verification fails or the executor falls over, the change undoes itself. A rollback that needs approval never auto-executes: it waits for you.
Escalates to a person
The incident is left as "needs a human", with what was tried and why it did not work. That is the honest ending when the catalog is not enough.
And if the diagnosis does not clear its rule's confidence threshold, Autopilot never applies anything: it leaves the plan as a suggestion.
05 — Postmortem · L6
By morning it is already written.
The last step fixes nothing: it tells the story. The timeline, what was tried, what resolved it and the evidence behind it, ready to paste into the channel or the ticket.
The timeline
From the first signal to the last pod going Ready, with real timestamps and what happened in between.
The action log
What ran, with what result and who approved it. In autonomous mode the approval is recorded as Autopilot's, not as somebody's.
The five whys
A draft analysis with proposed action items. The model writes it; you sign it.
06 — What you install, and what you don't
One agent, one outbound channel, and nothing else.
This is what you put inside your cluster so it can do it, and what stays yours.
The access model
Traditional dashboards reach into your API server. KubeBolt inverts the direction: the agent dials out, never in, and everything else rests on that.
Traditional model
Lens · Kite · K9s · direct kubectl
Client
kubeconfig
inbound
API Server
(public)
- API server must be reachable from the client
- Credentials on every user machine
- Private clusters need VPN or jump host
- Audit lives in the cluster, not centralized
KubeBolt model
Outbound agent — cluster stays private
Control Plane
KubeBolt
outbound
gRPC
PRIVATE CLUSTER
Agent
ServiceAccount
- API server stays private, never exposed publicly
- Credentials only on the agent (Kubernetes ServiceAccount)
- Works on private GKE/EKS and isolated VPCs
- Audit centralized at the KubeBolt control plane
Where the metrics come from
The agent adapts to what you already run: it takes kubelet metrics or queries your Prometheus, without you rebuilding your observability.
PRIVATE CLUSTER
DaemonSet mode
kubelet + Hubble
Prom-read mode
reads your Prometheus
Integrations
OpenCost · Trivy · Kyverno · Falco · ArgoCD
one channel
VictoriaMetrics
time series
Both modes coexist, each in its own pod with its own buffer.
- DaemonSet mode, no prerequisites: one pod per node takes kubelet metrics and, if you run Cilium, Hubble flow events.
- Prom-read mode, without touching your config: for AMP, Azure Monitor or GMP that can't remote_write outward, KubeBolt queries your Prometheus and forwards the series.
- Integrations auto-detect: OpenCost, Trivy, Kyverno, Falco and ArgoCD need no declaration. If they're in the cluster, they're used.
How much is kept
The signal rests in VictoriaMetrics. How far back you can look depends on your plan, or on your disk if you run it yourself.
Retention is how far back you can query. The Insights Engine detectors run on live state, so they don't depend on it.
Self-hosted, you're in charge.
Retention is a value in your Helm chart. We impose no caps on data that lives on your disk.
What it takes, concretely.
Around 1 GB per 100 pods at 30 days. A 60-pod cluster fits under 1 GB. Cardinality rules, not days.
Cluster state isn't persisted.
Pods, deployments and other objects live in memory and rebuild from informers on reconnect.
07 — And the dashboard
Everything you'd do with kubectl. From the browser.
Over that same channel you operate the whole cluster, not just watch it: terminal, logs, filesystem, port-forward and resource editing, from the browser.
01
Interactive terminal on any pod
Exec sessions with full stdin/stdout/stderr. Multi-pod, multi-cluster, straight from the browser. No kubeconfig on the client, no VPN, no access to the API server.
kubectl exec via agent02
Filesystem inspection at runtime
Browse the filesystem of running pods: directories, files generated by the app, configuration loaded at boot, certificates, mounted volumes with real data.
More useful than inspecting the static image: you see what's happening now, not what shipped.
runtime fs · not image03
Multi-pod log reading
On-demand reading at the workload level, with cross-pod aggregation across replicas. Read when you need them, no waiting on your observability stack and no continuous stream.
logs on-demand04
Port-forward without exposure
Access to internal UIs (Grafana, ArgoCD, private tooling) without LoadBalancers or public Ingress. The tunnel goes through the agent, the services stay private.
tunneling via agent05
Full resource CRUD
Create, update, delete, scale, restart, rollout via UI or API, with inline validation and native CRDs. Deleting Namespaces, Nodes, PV, PVC and RBAC resources is blocked by design.
kubectl surface06
Multi-cluster by default
One agent per cluster, all visible from the same control plane. No kubeconfig juggling, no mental context-switching. The cluster is just another filter in the UI.
fleet-native08 — If you use GitOps Roadmap · in design
If you use GitOps, the fix goes to the repo.
A fix in the cluster is gone with the next deploy. Autopilot already verifies every fix and reverts it if things get worse; what's coming is for it to be written in your repo, where the truth lives.
Root cause
Kobi, in Copilot or in Autopilot, identifies the failure
Source
the repo is resolved from the ArgoCD CR
Pull Request
minimal diff on the source
Merge
your team reviews, your checks rule
KubeBolt doesn't touch the cluster: it writes where the truth lives.
- A PR is never opened blind: first it confirms that this file produces exactly that workload. If it can't, it stays a recommendation and explains why.
- Your process rules: branch protection is respected. Auto-merge means merge when your checks pass, never skip them.
09 — What KubeBolt actually is
Most tools watch. KubeBolt operates.
“KubeBolt isn't a dashboard with AI added on. It's an AI operations platform for Kubernetes.
Your cluster self-diagnoses through deterministic runbooks and AI agents built on the Claude Agent SDK, and resolves the incident on its own when it can. Without exposing your API server: the agent opens the outbound connection and the cluster stays private.
The UI is excellent, but the product is two pieces: the Core (API), home to the correlation engine, and the agent that brings it the facts without exposing anything.”
Leafar Maina · Founder, CLM Cloud Solutions
Get the release notes → How we handle your data: trust & privacy →