KubeBolt docs
GitHub

Troubleshooting

Symptom, cause, fix — for the failures that actually happen: empty dashboards, the limited-access banner, 503s, a lost admin password, and an agent that can't reach the backend.

Find your symptom. Each entry says what you see, why it happens, and what to do. Scanner-specific failures live on their own pages — Trivy, Kyverno, Falco.

The agent is connected but the dashboards are empty

Symptom. The cluster shows as connected, the agent pod is Running, no errors anywhere — and every metric panel renders blank.

Cause. A schema mismatch between the agent and the backend. They ship as independently versioned artifacts coupled by one contract: the metric and label names the agent emits and the backend queries. When those disagree, samples still flow into the metrics store and the queries simply match nothing. Nothing crashes, which is exactly why this is hard to spot.

Fix. Check the backend log:

kubectl logs -n kubebolt deployment/kubebolt-api --tail=200 | grep "agent below minimum"

The WARN carries agent_version and min_agent_version. Upgrade the agent past the floor and the next refetch repopulates the panels:

helm upgrade kubebolt-agent \
  oci://ghcr.io/clm-cloud-solutions/kubebolt/helm/kubebolt-agent \
  -n kubebolt-system --reuse-values

No output from that grep and still-empty panels means the mismatch is the other way round — a current agent against a backend older than 1.10. See Compatibility for the matrix.

On a fleet, this is per cluster. A backend serving a mix of current and stale agents shows empty dashboards only for the clusters that lag, and logs one WARN per legacy agent on every reconnect.

Amber banner: “Limited access — N of M resource types restricted by RBAC”

Symptom. A warning strip under the topbar, and resource lists that are missing whole types.

Cause. The permission probe ran and found the ServiceAccount cannot list/watch some resource types. On agent-connected clusters this almost always means the agent’s ClusterRole is narrower than what the dashboard reads — rbac.mode=metrics instead of reader, or a hand-edited role.

Fix. Check what was probed — restricted types are dimmed with a shield icon in the sidebar, and GET /api/v1/cluster/permissions returns the full probe result — then widen the agent’s RBAC tier:

helm upgrade kubebolt-agent \
  oci://ghcr.io/clm-cloud-solutions/kubebolt/helm/kubebolt-agent \
  -n kubebolt-system --reuse-values --set rbac.mode=reader

The three tiers are described in RBAC. A missing optional CRD — Cilium policies, Argo CD applications, cert-manager certificates — does not count toward this banner and produces a separate, neutral notice instead. The banner is dismissible; dismissing it does not change any permission.

503 “cluster not connected” right after a switch or a restart

Symptom. Requests fail with 503 cluster not connected for a few seconds after switching clusters or restarting the backend. (Kobi still answers in that window — about KubeBolt and Kubernetes — but its cluster tools have nothing to read until the connector is up.)

Cause. The connector is still warming. A cold connect builds informer caches for every watched resource type and does not serve reads until they sync. The budget is 45 seconds by default (KUBEBOLT_CACHE_SYNC_TIMEOUT_SECONDS, floored at 5, or Administration → System → General at runtime).

Fix. Wait. It is a startup state, not an error, and the warm connector pool makes later switches instant. If it never resolves, read the next entry.

”Cluster unreachable” on a cluster that is reachable

Symptom. The API reaches the API server, the permission probe completes — and roughly 25 seconds later:

connect timed out after 25s for context in-cluster — agent may be stuck

The UI says Cluster unreachable. The message names an agent even on the in-cluster path, where there is none.

Cause. One slow informer blocks all of them. Reproduced on OpenShift 4.20 / Kubernetes 1.33.6: GET /api/v1/events returned ~138 MiB and the Events informer never finished, while Nodes and Deployments were ready in ~1.8 s and still could not be served. The outer connect deadline (25 s) is also shorter than the inner cache-sync budget (45 s), so the sync is killed before it can spend what it was given.

Fix. Two workarounds, neither of them free:

Confirm which one you are in before changing anything:

kubectl get events -A --no-headers | wc -l

Detail and the pending fixes are in OpenShift, where this was first reproduced. It is not OpenShift-specific — any churning cluster can hit it.

The admin password is lost

Symptom. You cannot log in and nobody has the generated first-boot password.

Cause. The password is printed to the API log exactly once, on the boot that seeds the admin user. A later restart does not reprint it.

Fix. First try the Secret: on an in-cluster install, the generated first-boot password is also stored in kubebolt-admin-password in the release namespace (it does not track later password changes):

kubectl -n kubebolt get secret kubebolt-admin-password \
  -o jsonpath='{.data.password}' | base64 -d ; echo

If the password was changed since, reset it. The helm upgrade path is the simple one — the new pod resets the password on boot and then starts normally:

helm upgrade kubebolt \
  oci://ghcr.io/clm-cloud-solutions/kubebolt/helm/kubebolt \
  -n kubebolt --reuse-values --set auth.resetAdminPassword='NEW-PASSWORD'

Running --reset-admin-password with kubectl exec inside the live pod does not work: the running API holds the embedded database’s lock, so the reset can’t open it. Without Helm, scale deployment/kubebolt-api to 0 and run the reset as a one-shot Job on the same PVC (kubebolt-api --reset-admin-password=NEW-PASSWORD with KUBEBOLT_DATA_DIR=/data) — the exact Job is in the chart README. On the single binary or single container, stop the process and run kubebolt --reset-admin-password=NEW-PASSWORD against the same data directory. Minimum length is 8 characters.

Clear auth.resetAdminPassword on the next helm upgrade. While it is set, every new pod resets the password on boot and the value sits in your release values.

On a fresh install, set the password instead of hunting for it: --set auth.adminPassword=... (see Authentication). The first-boot log line is only there on the boot that seeded the admin user:

kubectl -n kubebolt logs deployment/kubebolt-api | grep "Generated admin password"

The agent cannot reach the backend

Symptom. The agent pod runs but the cluster never appears, or the Add cluster wizard watches for five minutes and gives up.

Causes and fixes, in the order worth checking:

CheckWhat it looks likeFix
backendUrl has a schemehelm install fails, or the agent never dialsIt is a gRPC dial target, host:port, no https://
Wrong portConnection refusedA self-hosted backend’s <release>-agent-ingest Service listens on 9090; behind a TLS gateway or load balancer it is usually 443. KubeBolt Cloud’s is agent.kubebolt.io:443
Egress blockedTimeouts in the agent logThe agent only ever dials outbound. Allow egress to the backend host and port
TLS off against a TLS endpointHandshake failures--set tls.enabled=true for any endpoint that terminates TLS (typically 443)
Token missing or wrongThe channel closes right after connectingThe Secret must exist in the agent’s namespace with the token under token, and auth.ingestToken.existingSecret must name it. A backend in enforced agent auth mode rejects agents with auth.mode=disabled
Token belongs to another clusterRegisters under the wrong cluster, or is rejectedIssue a token per cluster in Administration → Agents & Ingest → Agent Tokens

Start here:

kubectl logs -n kubebolt-system -l app.kubernetes.io/name=kubebolt-agent --tail=100

Then confirm the backend saw anything at all: Administration → Agents & Ingest → Activity shows samples per second, active series, and a heartbeat list of the agents currently connected. If your cluster is absent from that list and no samples are arriving, the traffic never reached the backend — that is a network or address problem, not an auth one.

Kobi’s panel is missing

The Kobi button (bottom right) and ⌘J appear only when Kobi Copilot is configured — an API key saved in Administration → AI (Kobi), or set with KUBEBOLT_AI_API_KEY at boot. A key from the environment is logged at startup: kubectl -n kubebolt logs deployment/kubebolt-api | grep -i "AI copilot". See Enabling Kobi. If the panel opens and Kobi answers but can’t read the cluster, that is the connector, not the copilot: read the 503 entry above.

Still stuck

Open an issue with the backend version, the agent version, and the log line you saw, at github.com/clm-cloud-solutions/kubebolt, or tell us at kubebolt.io/feedback.