Troubleshooting
Symptom, cause, fix — for the failures that actually happen: empty dashboards, the limited-access banner, 503s, a lost admin password, and an agent that can't reach the backend.
Find your symptom. Each entry says what you see, why it happens, and what to do. Scanner-specific failures live on their own pages — Trivy, Kyverno, Falco.
The agent is connected but the dashboards are empty
Symptom. The cluster shows as connected, the agent pod is Running, no
errors anywhere — and every metric panel renders blank.
Cause. A schema mismatch between the agent and the backend. They ship as independently versioned artifacts coupled by one contract: the metric and label names the agent emits and the backend queries. When those disagree, samples still flow into the metrics store and the queries simply match nothing. Nothing crashes, which is exactly why this is hard to spot.
Fix. Check the backend log:
kubectl logs -n kubebolt deployment/kubebolt-api --tail=200 | grep "agent below minimum"
The WARN carries agent_version and min_agent_version. Upgrade the agent
past the floor and the next refetch repopulates the panels:
helm upgrade kubebolt-agent \
oci://ghcr.io/clm-cloud-solutions/kubebolt/helm/kubebolt-agent \
-n kubebolt-system --reuse-values
No output from that grep and still-empty panels means the mismatch is the
other way round — a current agent against a backend older than 1.10. See
Compatibility for the matrix.
On a fleet, this is per cluster. A backend serving a mix of current and stale agents shows empty dashboards only for the clusters that lag, and logs one WARN per legacy agent on every reconnect.
Amber banner: “Limited access — N of M resource types restricted by RBAC”
Symptom. A warning strip under the topbar, and resource lists that are missing whole types.
Cause. The permission probe ran and found the ServiceAccount cannot
list/watch some resource types. On agent-connected clusters this almost
always means the agent’s ClusterRole is narrower than what the dashboard reads
— rbac.mode=metrics instead of reader, or a hand-edited role.
Fix. Check what was probed — restricted types are dimmed with a shield
icon in the sidebar, and GET /api/v1/cluster/permissions returns the full
probe result — then widen the agent’s RBAC tier:
helm upgrade kubebolt-agent \
oci://ghcr.io/clm-cloud-solutions/kubebolt/helm/kubebolt-agent \
-n kubebolt-system --reuse-values --set rbac.mode=reader
The three tiers are described in RBAC. A missing optional CRD — Cilium policies, Argo CD applications, cert-manager certificates — does not count toward this banner and produces a separate, neutral notice instead. The banner is dismissible; dismissing it does not change any permission.
503 “cluster not connected” right after a switch or a restart
Symptom. Requests fail with 503 cluster not connected for a few seconds
after switching clusters or restarting the backend. (Kobi still answers in
that window — about KubeBolt and Kubernetes — but its cluster tools have
nothing to read until the connector is up.)
Cause. The connector is still warming. A cold connect builds informer caches
for every watched resource type and does not serve reads until they sync. The
budget is 45 seconds by default (KUBEBOLT_CACHE_SYNC_TIMEOUT_SECONDS, floored
at 5, or Administration → System → General at runtime).
Fix. Wait. It is a startup state, not an error, and the warm connector pool makes later switches instant. If it never resolves, read the next entry.
”Cluster unreachable” on a cluster that is reachable
Symptom. The API reaches the API server, the permission probe completes — and roughly 25 seconds later:
connect timed out after 25s for context in-cluster — agent may be stuck
The UI says Cluster unreachable. The message names an agent even on the in-cluster path, where there is none.
Cause. One slow informer blocks all of them. Reproduced on OpenShift 4.20 /
Kubernetes 1.33.6: GET /api/v1/events returned ~138 MiB and the Events
informer never finished, while Nodes and Deployments were ready in ~1.8 s and
still could not be served. The outer connect deadline (25 s) is also shorter
than the inner cache-sync budget (45 s), so the sync is killed before it can
spend what it was given.
Fix. Two workarounds, neither of them free:
- Drop
eventsfrom theget/list/watchrule of KubeBolt’s ClusterRole. The cluster comes up immediately; the Events tab and any insight that reads Events go empty, and the permission probe honestly reports one fewer resource type. - Raise the connect timeout —
KUBEBOLT_AGENT_PROXY_CONNECT_TIMEOUT(default25s), or the same value in Administration → Agents & Ingest → Configuration. It works, but it loosens a protection that exists for genuinely wedged agent-proxy clusters.
Confirm which one you are in before changing anything:
kubectl get events -A --no-headers | wc -l
Detail and the pending fixes are in OpenShift, where this was first reproduced. It is not OpenShift-specific — any churning cluster can hit it.
The admin password is lost
Symptom. You cannot log in and nobody has the generated first-boot password.
Cause. The password is printed to the API log exactly once, on the boot that seeds the admin user. A later restart does not reprint it.
Fix. First try the Secret: on an in-cluster install, the generated
first-boot password is also stored in kubebolt-admin-password in the release
namespace (it does not track later password changes):
kubectl -n kubebolt get secret kubebolt-admin-password \
-o jsonpath='{.data.password}' | base64 -d ; echo
If the password was changed since, reset it. The helm upgrade path is the
simple one — the new pod resets the password on boot and then starts normally:
helm upgrade kubebolt \
oci://ghcr.io/clm-cloud-solutions/kubebolt/helm/kubebolt \
-n kubebolt --reuse-values --set auth.resetAdminPassword='NEW-PASSWORD'
Running --reset-admin-password with kubectl exec inside the live pod does
not work: the running API holds the embedded database’s lock, so the reset
can’t open it. Without Helm, scale deployment/kubebolt-api to 0 and run the
reset as a one-shot Job on the same PVC (kubebolt-api --reset-admin-password=NEW-PASSWORD with KUBEBOLT_DATA_DIR=/data) — the
exact Job is in the chart README.
On the single binary or single container, stop the process and run
kubebolt --reset-admin-password=NEW-PASSWORD against the same data
directory. Minimum length is 8 characters.
Clear auth.resetAdminPassword on the next helm upgrade. While it is set,
every new pod resets the password on boot and the value sits in your release
values.
On a fresh install, set the password instead of hunting for it:
--set auth.adminPassword=... (see Authentication). The first-boot
log line is only there on the boot that seeded the admin user:
kubectl -n kubebolt logs deployment/kubebolt-api | grep "Generated admin password"
The agent cannot reach the backend
Symptom. The agent pod runs but the cluster never appears, or the Add cluster wizard watches for five minutes and gives up.
Causes and fixes, in the order worth checking:
| Check | What it looks like | Fix |
|---|---|---|
backendUrl has a scheme | helm install fails, or the agent never dials | It is a gRPC dial target, host:port, no https:// |
| Wrong port | Connection refused | A self-hosted backend’s <release>-agent-ingest Service listens on 9090; behind a TLS gateway or load balancer it is usually 443. KubeBolt Cloud’s is agent.kubebolt.io:443 |
| Egress blocked | Timeouts in the agent log | The agent only ever dials outbound. Allow egress to the backend host and port |
| TLS off against a TLS endpoint | Handshake failures | --set tls.enabled=true for any endpoint that terminates TLS (typically 443) |
| Token missing or wrong | The channel closes right after connecting | The Secret must exist in the agent’s namespace with the token under token, and auth.ingestToken.existingSecret must name it. A backend in enforced agent auth mode rejects agents with auth.mode=disabled |
| Token belongs to another cluster | Registers under the wrong cluster, or is rejected | Issue a token per cluster in Administration → Agents & Ingest → Agent Tokens |
Start here:
kubectl logs -n kubebolt-system -l app.kubernetes.io/name=kubebolt-agent --tail=100
Then confirm the backend saw anything at all: Administration → Agents & Ingest → Activity shows samples per second, active series, and a heartbeat list of the agents currently connected. If your cluster is absent from that list and no samples are arriving, the traffic never reached the backend — that is a network or address problem, not an auth one.
Kobi’s panel is missing
The Kobi button (bottom right) and ⌘J appear only when Kobi Copilot is
configured — an API key saved in Administration → AI (Kobi), or set with
KUBEBOLT_AI_API_KEY at boot. A key from the environment is logged at startup:
kubectl -n kubebolt logs deployment/kubebolt-api | grep -i "AI copilot". See
Enabling Kobi. If the panel opens and Kobi answers but
can’t read the cluster, that is the connector, not the copilot: read the 503
entry above.
Still stuck
Open an issue with the backend version, the agent version, and the log line you saw, at github.com/clm-cloud-solutions/kubebolt, or tell us at kubebolt.io/feedback.