Non-root OpenShift containers can lose capabilities after SCC admission
Red Hat’s runtime trace shows why a capability accepted by an SCC and requested in a PodSpec may still disappear when the container process changes UID.
A security context constraint can admit an OpenShift pod, and its PodSpec can explicitly request a Linux capability, without that capability surviving into a non-root process. A new Red Hat Developer engineering analysis traces that failure across SCC admission, CRI-O’s OCI-spec construction and the kernel’s UID-transition rules.
That distinction matters for platform teams debugging workloads that appear correctly configured. In Red Hat’s example, a monitoring sidecar running as UID 1000 requests CAP_NET_RAW so it can perform ICMP health checks. The SCC allows the capability, the pod is admitted and the container starts, but ping fails because the process has an empty effective capability set.
Admission is not runtime enforcement
The core lesson is that SCC answers whether a pod’s requested security configuration is allowed into the cluster. It does not grant privileges directly to a running process. After admission, CRI-O translates the container security context into an OCI runtime specification, and the Linux kernel determines which capabilities survive process setup.
For a non-privileged container, CRI-O places requested capabilities into the bounding, permitted and effective sets. It does not populate the ambient set. When the runtime changes the process from UID 0 to a non-zero UID, the kernel clears the permitted and effective sets. The bounding set can still show CAP_NET_RAW, while the effective set—the one the kernel checks when ping opens a raw socket—is empty.
That behavior is deliberate rather than an SCC or CRI-O defect. The post says CRI-O omits ambient capabilities because Kubernetes expects a switch to a non-root user to drop capabilities. It also points to Kubernetes enhancement proposal 2763 as the unfinished path toward explicit ambient-capability support.
The smallest fix changes the image trust boundary
Red Hat presents file capabilities as the most targeted workaround: set cap_net_raw=+ep on the required binary during the image build. When that binary executes, the kernel can promote the file capability into the process’s permitted and effective sets, provided the capability remains in the bounding set.
The fix is narrower than running the whole container as root or privileged, but it is not free. Extended attributes must survive the image build and storage path, and the image itself now carries privileged metadata. Platform teams therefore need image-provenance controls that can establish who added the file capability and whether the artifact was changed afterward.
Running as root avoids the UID transition, but conflicts with the production-default restricted-v2 SCC. Privileged mode is broader still: according to Red Hat’s trace, it fills all five capability sets and bypasses the normal add/drop calculation, while also disabling SELinux, AppArmor and seccomp protections.
A better debugging order
The practical change is to stop treating successful SCC admission as proof of runtime privilege. After confirming SCC and PodSpec settings, operators should inspect /proc/1/status inside the container and compare the bounding, permitted, effective, inheritable and ambient sets. If the bounding set contains the requested capability but the effective set is empty for a non-root process, the UID transition—not SCC selection—is the likely boundary to investigate.
That makes file capabilities a specific, auditable remediation and keeps root or privileged execution as explicit risk decisions rather than reflexive fixes.
sources
- Why your non-root container dropped its capabilities (and how to fix it)developers.redhat.com
comments · 0