How it works · the incident loop

It is 3am and something breaks. Here is what happens next.

The same incident the home page resolves in 87 seconds, step by step: what detects it, what the model reads, what it may touch and what is written down by morning.

01 — Detect · L1

Nobody pages you. It already noticed.

Twenty-four deterministic rules watch the cluster's state: crash loops, OOM kills, images that will not pull, certificates about to expire, nodes going away. Zero configuration, and not one AI credit. And they do not hand you a bare finding: each one arrives with its recommendation, so you can fix it yourself even when that incident is not in the catalog Autopilot handles on its own.

The six layers an incident travels through

L1

Detector

rules over events

100% deterministic
L2

Router

classifies and prioritizes

AI
L3

Investigator

correlates logs, metrics and deploys

AI
L4

Planner

builds the plan

AI
L5

Executor

executes behind a deterministic guardrail

AI + deterministic
L6

Postmortem

timeline and evidence

AI

Each chapter below is one of them. Only the first is 100% deterministic; the last one that touches production works behind a guardrail that is deterministic too.

L1 Insights Engine (L1): 24 zero-config rules detect the incident before any model gets involved. See the 24 rules →

  • The rules run on current state, not on history, so they work from the first minute and do not depend on how much retention your plan has.
  • Detecting costs no credits. They are spent only when something needs investigating, and the next layer makes that call.
  • Every rule is a Skill: it carries its own confidence threshold and says which catalog actions are valid for that incident.

02 — Diagnose · L2 to L4

The model proposes. It never executes.

When that signal becomes a question, Kobi reasons and proposes the answer, but doesn't execute it: you authorize, and it's all audited.

The LLM has no route to the cluster
can't execute gets no credentials doesn't evaluate RBAC

The tools run in the backend. The model only sees results.

  • BYOK (your own key or endpoint: Azure OpenAI, vLLM, Ollama) applies to OSS and EE Self-Hosted: AI traffic goes from your install to your provider, never through KubeBolt. On SaaS the AI is managed.
  • You authorize each action and it runs with the agent's RBAC; KubeBolt never impersonates your credentials or your Kubernetes RBAC. Who gets read-only or edit is set by your org admin, and every action is audited: who authorized it, when, and which resources it touched.
  • A tampered log can't execute anything. Logs are text a third party may have written; that's why the model proposes instead of executing.

The rule says what broke; why it broke needs context: events, logs, metrics and recent changes. It is the only thing that crosses the channel.

Leaves
Kubelet metrics L3/L4/L7 network flows Kubernetes objects Logs (on demand)
Never crosses
Secret values Network payloads Your kubeconfig Another org's data

The agent opens the connection. The cluster exposes no port.

  • Secret values are redacted before they leave the API layer — not a prompt policy, a server-side check.
  • Metadata, not content. Each flow carries source, destination, verdict, status code and latency — never the request body.
  • Logs aren't shipped in bulk. They're read only when you ask for a diagnosis that needs them.
  • Self-hosted, nothing reaches KubeBolt: no product telemetry, no usage ping.

03 — Remediate · L5

Kobi acts alone. You approve anything destructive.

With nobody to ask, the signal fires Kobi: it detects, investigates, remediates and documents. The model proposes; a deterministic guardrail checks the action against the catalog before anything touches production.

  • live

    Anything destructive is approved by a person

    A rollback always requires approval, at any autonomy level.

  • live

    If verification fails, Autopilot reverts

    The change goes back to its previous state, automatically.

  • live

    Every action is logged

    With its result and who approved it.

Three autonomy levels: suggest only (the default), approve and execute, autonomous. And off, if you want. With a confidence gate and per-cluster blocked namespaces.

Model planL4 Action Registrywhitelist in code Executorapplies the validated

If a step isn't on the list, the plan is rejected and escalates to a person.

  • The plan is validated against code, not against a prompt. A new action needs a release, not a different instruction.
  • An unattended window (days, hours, timezone) and a platform-level autonomy ceiling: outside the window, the incident is recorded, not acted on.

04 — Verify and revert

If the fix does not hold, it undoes it and calls you.

Applying is not resolving. Autopilot captures the previous values before touching anything, checks the result and, if it does not pass, restores and escalates to a person. It ships today; it is not on the roadmap.

01live

Captures the previous state

The values it is about to change are captured before anything is applied, so the revert restores exactly what was there.

02live

Checks the result

It waits for pods to go Ready and to stop restarting. An action applying without error is not enough.

03live

Reverts if it fails

In autonomous mode, if verification fails or the executor falls over, the change undoes itself. A rollback that needs approval never auto-executes: it waits for you.

04live

Escalates to a person

The incident is left as "needs a human", with what was tried and why it did not work. That is the honest ending when the catalog is not enough.

And if the diagnosis does not clear its rule's confidence threshold, Autopilot never applies anything: it leaves the plan as a suggestion.

05 — Postmortem · L6

By morning it is already written.

The last step fixes nothing: it tells the story. The timeline, what was tried, what resolved it and the evidence behind it, ready to paste into the channel or the ticket.

The timeline

From the first signal to the last pod going Ready, with real timestamps and what happened in between.

The action log

What ran, with what result and who approved it. In autonomous mode the approval is recorded as Autopilot's, not as somebody's.

The five whys

A draft analysis with proposed action items. The model writes it; you sign it.

06 — What you install, and what you don't

One agent, one outbound channel, and nothing else.

This is what you put inside your cluster so it can do it, and what stays yours.

The access model

Traditional dashboards reach into your API server. KubeBolt inverts the direction: the agent dials out, never in, and everything else rests on that.

Traditional model

Lens · Kite · K9s · direct kubectl

  • API server must be reachable from the client
  • Credentials on every user machine
  • Private clusters need VPN or jump host
  • Audit lives in the cluster, not centralized

KubeBolt model

Outbound agent — cluster stays private

  • API server stays private, never exposed publicly
  • Credentials only on the agent (Kubernetes ServiceAccount)
  • Works on private GKE/EKS and isolated VPCs
  • Audit centralized at the KubeBolt control plane

Where the metrics come from

The agent adapts to what you already run: it takes kubelet metrics or queries your Prometheus, without you rebuilding your observability.

Both modes coexist, each in its own pod with its own buffer.

  • DaemonSet mode, no prerequisites: one pod per node takes kubelet metrics and, if you run Cilium, Hubble flow events.
  • Prom-read mode, without touching your config: for AMP, Azure Monitor or GMP that can't remote_write outward, KubeBolt queries your Prometheus and forwards the series.
  • Integrations auto-detect: OpenCost, Trivy, Kyverno, Falco and ArgoCD need no declaration. If they're in the cluster, they're used.

How much is kept

The signal rests in VictoriaMetrics. How far back you can look depends on your plan, or on your disk if you run it yourself.

OSS · self-hosted∞
Free15 days
Team30 days
Business90 days
Enterprise365 days

Retention is how far back you can query. The Insights Engine detectors run on live state, so they don't depend on it.

Self-hosted, you're in charge.

Retention is a value in your Helm chart. We impose no caps on data that lives on your disk.

What it takes, concretely.

Around 1 GB per 100 pods at 30 days. A 60-pod cluster fits under 1 GB. Cardinality rules, not days.

Cluster state isn't persisted.

Pods, deployments and other objects live in memory and rebuild from informers on reconnect.

07 — And the dashboard

Everything you'd do with kubectl. From the browser.

Over that same channel you operate the whole cluster, not just watch it: terminal, logs, filesystem, port-forward and resource editing, from the browser.

01

Interactive terminal on any pod

Exec sessions with full stdin/stdout/stderr. Multi-pod, multi-cluster, straight from the browser. No kubeconfig on the client, no VPN, no access to the API server.

kubectl exec via agent

02

Filesystem inspection at runtime

Browse the filesystem of running pods: directories, files generated by the app, configuration loaded at boot, certificates, mounted volumes with real data.

More useful than inspecting the static image: you see what's happening now, not what shipped.

runtime fs · not image

03

Multi-pod log reading

On-demand reading at the workload level, with cross-pod aggregation across replicas. Read when you need them, no waiting on your observability stack and no continuous stream.

logs on-demand

04

Port-forward without exposure

Access to internal UIs (Grafana, ArgoCD, private tooling) without LoadBalancers or public Ingress. The tunnel goes through the agent, the services stay private.

tunneling via agent

05

Full resource CRUD

Create, update, delete, scale, restart, rollout via UI or API, with inline validation and native CRDs. Deleting Namespaces, Nodes, PV, PVC and RBAC resources is blocked by design.

kubectl surface

06

Multi-cluster by default

One agent per cluster, all visible from the same control plane. No kubeconfig juggling, no mental context-switching. The cluster is just another filter in the UI.

fleet-native

08 — If you use GitOps Roadmap · in design

If you use GitOps, the fix goes to the repo.

A fix in the cluster is gone with the next deploy. Autopilot already verifies every fix and reverts it if things get worse; what's coming is for it to be written in your repo, where the truth lives.

01

Root cause

Kobi, in Copilot or in Autopilot, identifies the failure

02

Source

the repo is resolved from the ArgoCD CR

03

Pull Request

minimal diff on the source

04

Merge

your team reviews, your checks rule

KubeBoltopens the PR Your repoGitHub · GitLab ArgoCDsyncs Your clusterno drift

KubeBolt doesn't touch the cluster: it writes where the truth lives.

  • A PR is never opened blind: first it confirms that this file produces exactly that workload. If it can't, it stays a recommendation and explains why.
  • Your process rules: branch protection is respected. Auto-merge means merge when your checks pass, never skip them.

09 — What KubeBolt actually is

Most tools watch. KubeBolt operates.

“KubeBolt isn't a dashboard with AI added on. It's an AI operations platform for Kubernetes.

Your cluster self-diagnoses through deterministic runbooks and AI agents built on the Claude Agent SDK, and resolves the incident on its own when it can. Without exposing your API server: the agent opens the outbound connection and the cluster stays private.

The UI is excellent, but the product is two pieces: the Core (API), home to the correlation engine, and the agent that brings it the facts without exposing anything.”

Leafar Maina · Founder, CLM Cloud Solutions