KubeBolt docs
GitHub

Compatibility

Which agent version pairs with which backend, what breaks when they disagree, and which Kubernetes versions and providers are supported.

KubeBolt and kubebolt-agent ship as independently versioned artifacts — the agent has its own chart, image and release tag. They are coupled through exactly one contract: the metric and label schema the agent emits and the backend’s queries consume.

Current: KubeBolt 2.3.0, agent 1.4.2.

Backend and agent

KubeBoltAgentSchemaNotes
1.10.x – 2.3.x1.0.x – 1.4.xv1.0, Prometheus-canonicalCurrent. Any 1.x agent from 1.0 pairs with any backend from 1.10. Nothing the agent depends on changed from 2.0 through 2.3 — the same metric and label schema, and agents 1.4.1 and 1.4.2 change no flag, value or protocol
1.10.x and later0.2.xmismatchThe agent emits the legacy schema; the backend’s queries look for canonical names. Empty dashboards
1.9.x and earlier1.0.xmismatchThe reverse. Empty dashboards. Do not run this combination
1.9.x and earlier0.2.xlegacy v0.xBoth sides pre-canonical. Works, but the newer features (right-sizing P95, network drops, external endpoints with FQDN) only exist on the canonical pair
1.5.x0.1.xearlyPre-3-tier-RBAC. Stable, but no agent-proxy and no Hubble flows

So the rule is short: agent ≥ 1.0 with backend ≥ 1.10, and keep them on the same generation. Agent 1.1 through 1.4 are reliability, cardinality and OpenCost-sourcing releases on the same schema — 1.4.0 widens the default dropNetworkInterfaces list, 1.4.1 is a security patch, and 1.4.2 makes the chart install on OpenShift out of the box — so an agent upgrade inside the 1.x line is never gated on a backend upgrade.

Run agent chart 1.4.2 with agent image 1.4.2, which is the default. The 1.4.2 image declares a numeric user; on OpenShift, chart 1.4.2 with an older image pinned through image.tag fails the runAsNonRoot check.

A mismatch does not crash anything. Both sides start, the agent registers, samples reach the metrics store. The only visible symptom is that the panels for that cluster render empty — which is why it gets misdiagnosed as a data problem. See Troubleshooting.

Detecting a mismatch

A legacy agent connecting to a current backend logs, once per registration:

{
  "level": "WARN",
  "msg": "agent below minimum version — legacy schema",
  "agent_version": "0.2.2",
  "min_agent_version": "1.0.0",
  "hint": "upgrade kubebolt-agent helm chart to >=1.0.0; v0.x emits the legacy schema and dashboards will render empty"
}

The check is fail-soft: the agent still connects and still ships samples. An empty or unparseable agent_version is silent, because not every client sets it.

Upgrading both

Order is operationally forgiving, but roll the backend first — the agent reconnects at registration:

# --reset-then-reuse-values (Helm 3.13+) applies the new chart defaults under
# your existing overrides; plain --reuse-values can fail to render when the
# chart gained value blocks since your last install.
helm upgrade kubebolt \
  oci://ghcr.io/clm-cloud-solutions/kubebolt/helm/kubebolt \
  --version 2.3.0 -n kubebolt --reset-then-reuse-values

helm upgrade kubebolt-agent \
  oci://ghcr.io/clm-cloud-solutions/kubebolt/helm/kubebolt-agent \
  --version 1.4.2 -n kubebolt-system --reset-then-reuse-values

# Verify: no output means no legacy agent is left
kubectl logs -n kubebolt deployment/kubebolt-api --tail=50 | grep "agent below minimum"

If you pin api.resources in your own values file, check its memory limit: from 2.3.0 the chart’s default is 1Gi (request 128Mi, was 256Mi / 64Mi), because the API caches every Events object it watches and a busy cluster outgrew the old limit before its first sync finished.

On a fleet, upgrade every agent before declaring the rollout done. A partial rollout is recoverable — finish it and the next refetch repopulates the dashboards — but until then the lagging clusters show empty panels and the backend logs one WARN each on every reconnect.

Avoid plain --reuse-values when the chart’s own defaults have moved; it pins you to the values you installed with. Prefer --reset-then-reuse-values or an explicit values file.

What changed at agent 1.0

Agent 1.0.0 renamed every metric and label to follow the Prometheus convention for Kubernetes — the de-facto schema of cAdvisor, kube-state-metrics, node-exporter and Hubble. There is no dual emission: the agent ships only the canonical names. The highlights:

Future major schema changes bump the agent major version and the backend major version in lockstep. Minor and patch agent releases stay backward-compatible with the current backend minor: additive metrics and labels only, never renames.

Kubernetes versions

VersionSupport
1.24 and newerFull. EndpointSlice is GA and CronJob is v1, which the insights engine relies on
1.20 – 1.23Works with feature degradation. Older clusters fall back to v1.Endpoints, which KubeBolt does not watch, so the service-without-endpoints insight degrades

Check yours:

kubectl version -o json | jq -r .serverVersion.gitVersion

Providers

ProviderMetrics ServerNotes
Amazon EKSNot installed by defaultInstall it (manifest or EKS add-on) — see the EKS guide. Managed Prometheus (AMP) supported, with SigV4 auth via IRSA or Pod Identity
Google GKEPre-installedManaged Prometheus (GMP) supported, with Workload Identity
Azure AKSPre-installedAzure Monitor managed Prometheus supported, with Workload Identity
OpenShiftProvided by the clusterInstalls under the default restricted-v2 SCC with no workarounds from KubeBolt 2.3.0 and agent 1.4.2. Older releases need the workarounds in the OpenShift guide, which also covers the one item still open
k3s / k0sPre-installedWorks out of the box
kubeadmManual installkubectl apply -f the metrics-server release manifest
Docker DesktopManual installSee the Docker Desktop guide
MinikubeAddonminikube addons enable metrics-server
kind / k3dManual installMay need --kubelet-insecure-tls. Note the Falco caveat: the “nodes” share one kernel

The 1.10 release was validated across GKE with Dataplane V2, EKS, AKS, and GKE with Calico OSS.

Metrics Server is recommended, not required. Without it the Overview’s live commitment bars read “no data”; everything else — including historical CPU and memory through the agent — keeps working.

Release notes

Per-version notes are published on GitHub releases, and the agent keeps its own changelog in the same repository. See Changelog.