OpenShift
From KubeBolt 2.3.0 and agent 1.4.2, KubeBolt installs under OpenShift's default restricted-v2 SCC with no workarounds. What changed, what is still open, and the workarounds for older releases.
Status — fixed in KubeBolt 2.3.0 and agent 1.4.2, pending re-verification
on OpenShift. The problems were reported on KubeBolt 1.23.1 / OpenShift 4.20 /
Kubernetes 1.33.6 (2026-08-16): the web container did not start under the
default restricted-v2 SCC, the agent chart pinned a UID the SCC rejects, and
on a busy cluster the in-cluster connection timed out and blamed an agent that
does not exist. One part stays open — see Still open. On
releases before 2.3.0 / agent 1.4.2, use the
workarounds.
Installing on 2.3.0 and later
A standard install works under restricted-v2 with no SCC grant and no patch.
What changed to make that true:
- The
webimage runs under any UID in group 0 — whatrestricted-v2assigns. Every path nginx and its entrypoint write to belongs to group root and is group-writable. As root it behaves as before, so the same image keeps working on vanilla Kubernetes and Docker Compose. - The
webDeployment runs under the chart’s ServiceAccount, so an SCC grant for the release reaches it, and it does not mount its token: nginx never calls the Kubernetes API. api.podSecurityContext/api.securityContext, and the same forweb, are available and empty by default. On OpenShift leaverunAsUserunset:restricted-v2assigns a UID from the namespace’s range and rejects a pod that asks for one outside it.- In-cluster and kubeconfig connects are no longer cut at 25 s. That
deadline is an agent-proxy setting, written for a wedged agent; it now applies
to agent-proxy clusters only. Direct connections are bounded by Cluster
connect timeout (Administration → System → General, or
KUBEBOLT_CACHE_SYNC_TIMEOUT_SECONDS; 45 s by default), and “agent may be stuck” is no longer shown where there is no agent.
Size the API to its Events
The API caches every resource it watches, Events included. On a test cluster with 161 MB of Events the API held ~253 MiB once synced — right at the old default limit of 256Mi, so it never finished its first sync and its liveness probe restarted it every ~80 s. From 2.3.0 the default API memory limit is 1Gi (request 128Mi). On a cluster whose Events are larger still, raise it:
helm upgrade kubebolt \
oci://ghcr.io/clm-cloud-solutions/kubebolt/helm/kubebolt \
--version 2.3.0 -n kubebolt --reset-then-reuse-values \
--set api.resources.limits.memory=2Gi
If you pin api.resources in your own values file, check its memory limit when
you upgrade: your value wins over the new default.
The agent
From agent 1.4.2 the chart detects OpenShift — a cluster that serves
security.openshift.io/v1 — and drops its default runAsUser: 65532, leaving
the UID to the SCC. The image’s USER is now numeric (65532:65532, the same
UID as before), so runAsNonRoot: true still verifies. A runAsUser you set
yourself is kept.
Two cases need a hand:
- Rendering without cluster access (
helm template, some GitOps tools): the chart cannot see the API group, so pass--api-versions security.openshift.io/v1, or set--set podSecurityContext.runAsUser=nullas before. - Run chart 1.4.2 with image 1.4.2 (the default). Chart 1.4.2 on OpenShift
with an older image pinned through
image.tagfails therunAsNonRootcheck, because the older image’s user is not numeric.
Nothing else in the agent needs elevated privileges: the DaemonSet declares no
hostPath, hostNetwork or hostPID, and readOnlyRootFilesystem: true is
compatible with restricted-v2 as-is.
Still open
One slow informer still holds the others back. The connector waits on the whole shared informer factory, so on a cluster with a very large Events collection, Nodes and Pods that were ready in under two seconds are not served until Events finish syncing. Since 2.3.0 that wait runs up to the Cluster connect timeout instead of being cut at 25 s, so the cluster does come up — it just comes up as late as its slowest informer. If Events on your cluster take longer than the timeout, raise Cluster connect timeout rather than the agent-proxy setting.
The fix being tracked: bound the Events informer (a limit, a field selector or
a shorter window), sync it asynchronously, or wait per informer so a slow one
degrades its own tab instead of the whole context.
What was reported, and why it happened
The rest of this page documents the original problems. It explains the workarounds below and is useful if you run an older release.
The web container crash-looped
A standard install left the web container in CrashLoopBackOff while
everything else ran:
NAME READY STATUS RESTARTS
kubebolt-api-... 1/1 Running 0
kubebolt-victoriametrics-0 1/1 Running 0
kubebolt-web-... 0/1 CrashLoopBackOff 7
/docker-entrypoint.sh: Launching /docker-entrypoint.d/40-api-backend.sh
sed: can't create temp file '/etc/nginx/conf.d/default.confXXXXXX': Permission denied
restricted-v2 assigns each pod a random UID from the namespace’s range
(openshift.io/sa.scc.uid-range) and always puts it in GID 0. The web
image is built on nginx:1.27-alpine, where every path nginx writes to was
root:root 0755 — group-readable, not group-writable. Four paths were denied:
| Path | Written by |
|---|---|
/etc/nginx/conf.d | the entrypoint’s sed -i, which substitutes API_BACKEND |
/var/cache/nginx | nginx master at startup (client_temp and friends) |
/var/log/nginx | nginx master, before it drops privileges |
/run | the PID file |
Two plausible theories the evidence ruled out:
- It was not the volumes. PersistentVolumeClaims are unaffected:
restricted-v2supplies anfsGroupfrom the namespace range and the kubelet chowns the volume to it. The VictoriaMetrics PVC and the API’s/dataPVC worked untouched, which is why those two pods wereRunning. - A
securityContextin the chart could not fix it, and addingrunAsUseractively breaks it:restricted-v2uses theMustRunAsRangestrategy, so a pod requesting a UID outside the namespace’s range is rejected at admission. The fix belonged in the image, which is where 2.3.0 put it.
The in-cluster context timed out
Independent of the SCC problem. The API detected its in-cluster config, reached the API server and finished the permission probe — then, about 25 seconds later:
connect timed out after 25s for context in-cluster — agent may be stuck
The UI showed Cluster unreachable even though nothing was unreachable. The
reporter removed only events from the get/list/watch rule of the
kubebolt ClusterRole and restarted:
with events | without events | |
|---|---|---|
| Permission probe | 26/31 | 25/31 |
| Informer caches | never sync | synced ~1.8 s after the probe |
| UI | Cluster unreachable | Active — 5/5 nodes, 67/67 deployments, 77 namespaces |
On that cluster oc get events -A took ~32 s and GET /api/v1/events returned
~138 MiB, while Nodes, Pods and Deployments answered in under a second each.
Three things compounded:
- Events are unbounded in a way nothing else is. Everything else is bounded by what is deployed; Events are bounded by what happens.
- One slow informer blocks every other one — the part that is still open.
- The outer deadline was shorter than the inner budget. The agent-proxy connect deadline (25 s) also governed the direct in-cluster connect, while the cache-sync budget was 45 s, so a sync progressing normally toward 30 s was killed at 25. 2.3.0 removed that deadline from direct connections.
Before 2.3.0: workarounds
Only for releases before KubeBolt 2.3.0 / agent 1.4.2. Upgrading removes the need for all of them.
web: grant the anyuid SCC
Fastest path. Needs cluster-admin. The pod runs as root, exactly as it does
on vanilla Kubernetes.
oc adm policy add-scc-to-user anyuid -z default -n kubebolt
oc rollout restart deployment/kubebolt-web -n kubebolt
-z default is right for those releases: the web Deployment set no
serviceAccountName, so it ran under the namespace’s default ServiceAccount.
This weakens the namespace’s security posture; prefer the next option if your
platform team will not grant anyuid.
web: keep restricted-v2
An initContainer renders the nginx config into an emptyDir, the main
container bypasses the entrypoint scripts, and the three remaining write paths
become emptyDir mounts — writable because the SCC’s fsGroup applies to them.
oc patch deployment/kubebolt-web -n kubebolt --type=strategic -p '
spec:
template:
spec:
initContainers:
- name: render-config
image: ghcr.io/clm-cloud-solutions/kubebolt/web:2.2.1
command: ["/bin/sh", "-c"]
args:
- sed "s|api:8080|${API_BACKEND}|g" /etc/nginx/conf.d/default.conf > /mnt/conf/default.conf
env:
- name: API_BACKEND
value: kubebolt-api:8080
volumeMounts:
- name: nginx-conf
mountPath: /mnt/conf
containers:
- name: web
command: ["nginx", "-g", "daemon off;"]
volumeMounts:
- name: nginx-conf
mountPath: /etc/nginx/conf.d
- name: nginx-cache
mountPath: /var/cache/nginx
- name: nginx-log
mountPath: /var/log/nginx
- name: nginx-run
mountPath: /run
volumes:
- name: nginx-conf
emptyDir: {}
- name: nginx-cache
emptyDir: {}
- name: nginx-log
emptyDir: {}
- name: nginx-run
emptyDir: {}
'
Pin the initContainer image to the tag you installed, and set API_BACKEND to
<release-name>-api:8080.
A helm upgrade reverts this patch. Re-apply it after every upgrade, or
keep it in a post-render kustomization.
Agent 1.4.1 and earlier: drop the pinned UID
helm upgrade --install kubebolt-agent \
oci://ghcr.io/clm-cloud-solutions/kubebolt/helm/kubebolt-agent \
--namespace kubebolt \
--set podSecurityContext.runAsUser=null \
--set backendUrl=<host:port>
runAsNonRoot: true still holds, because the UID OpenShift assigns is never 0.
--set-json 'podSecurityContext={}' does not work — Helm merges maps, so an
empty map leaves the pinned UID in place. Only =null removes the key.
The in-cluster timeout
- Drop
eventsfrom the ClusterRole. The cluster comes up immediately; the cost is that the Events tab and any insight that reads Events go empty, and the probe honestly reports one fewer resource type. - Raise the agent-proxy connect timeout —
KUBEBOLT_AGENT_PROXY_CONNECT_TIMEOUT(default25s), or Administration → Agents & Ingest → Configuration. It works, but it is the wrong knob twice over: it is an agent-proxy resilience setting, and loosening it weakens the protection it exists for on clusters that do use agents.
Reproducing without an OpenShift cluster
Run the image the way restricted-v2 would — any UID works as long as it is
not in /etc/passwd and the GID is 0, which is the whole of what OpenShift
does differently:
docker run --rm --user 1000700000:0 --entrypoint /bin/sh \
ghcr.io/clm-cloud-solutions/kubebolt/web:2.3.0 -c '
sed -i "s|api:8080|x:1|g" /etc/nginx/conf.d/default.conf
ls -ld /etc/nginx/conf.d /var/cache/nginx /var/log/nginx /run
'
On 2.3.0 the sed succeeds and the four paths are owned by group 0 and
group-writable. On 2.2.1 and earlier the same command fails with
Permission denied.
Fixes and where they shipped
| Fix | Since |
|---|---|
web image tolerates an arbitrary UID (chgrp 0 + chmod g=u on the four paths) | 2.3.0 |
web Deployment runs under the chart’s ServiceAccount, token not mounted | 2.3.0 |
Optional api / web podSecurityContext and securityContext, empty by default | 2.3.0 |
| The agent-proxy connect deadline no longer applies to in-cluster / kubeconfig connects, and “agent may be stuck” is not shown where there is no agent | 2.3.0 |
| Default API memory limit 1Gi (request 128Mi), was 256Mi | 2.3.0 |
The agent chart drops its default runAsUser on OpenShift; the agent image’s USER is numeric | agent 1.4.2 |
| One slow informer no longer holds the others back | open |
References
- Managing security context constraints — the
MustRunAsRangestrategy and the namespace UID range - Pod Security Standards in OpenShift 4.11+