Part 2: The seccomp profile for kubernetes pod nobody writes by hand
Your container makes about 70 of Linux's 300-odd syscalls. Here's how kguardian works out which 70, and how it tells you when that answer goes stale. Part 2 of seven. Parts 1 to 3 are the operator's tour, 4 to 7 open t
Your container makes about 70 of Linux's 300-odd syscalls. Here's how kguardian works out which 70, and how it tells you when that answer goes stale.
Part 2 of seven. Parts 1 to 3 are the operator's tour, 4 to 7 open the hood on the eBPF. Part 1 covered the network side. Syscalls are harder, in an interesting way.
Why seccomp is worse than NetworkPolicy
A wrong NetworkPolicy degrades. A connection fails, something retries, you get a log line you can act on.
A wrong seccomp profile doesn't degrade. It's an allowlist evaluated in the kernel on every syscall, so SCMP_ACT_ERRNO on a syscall your runtime needed returns EPERM into a code path that has never seen EPERM. SCMP_ACT_KILL_PROCESS is blunter still. The failure is immediate, total, and usually reported as something unrelated.
So almost nobody ships one. You stay on RuntimeDefault, which is generic enough to cover every container ever built, or you strace it in staging and get a profile that's correct for whatever ran on a Tuesday.
Same premise as part 1: don't guess. The kernel already knows which syscalls a process makes. It's the one servicing them.
Capture: five tiers, one of them safe
kguardian controller traces syscalls per pod with eBPF and keeps a union of the names it saw. There are five tiers, set with syscalls.captureLevel (default full):
| Tier | Traces | Safe to enforce? |
|---|---|---|
full |
every syscall | yes |
high |
all but hot-path noise (read, write, futex, mmap, โฆ) |
no |
medium |
low plus network, file-permission and process-lifecycle families |
no |
low |
57 security-relevant syscalls (execve, ptrace, mount, bpf, โฆ) |
no |
custom |
your list | no |
That last column is the trap the whole feature is built around. A profile is an allowlist, so a partial capture doesn't give you a slightly worse profile. It gives you one that denies exactly the syscalls the tier never traced. A low capture sees ptrace and misses read. Ship that as SCMP_ACT_ERRNO and the container can't read a file.
The lower tiers are for clusters that want the security signal and will never enforce. If you do intend to ship profiles, full isn't the expensive choice you'd assume: it costs about what low costs, because syscall names are de-duplicated inside BPF per pod, so the hundred-thousandth read never crosses into userspace. That's part 6, and it's what makes this affordable.
Then let it run under real load for a full business cycle. For a CronJob, several runs.

Seccomp Profiles: one workload observed at full capture, no CR deployed yet
Completeness is checked across every pod that ever contributed, so one historical low run keeps a workload incomplete until the union refills. It then follows you everywhere: the Capture badge above, capture.complete: false in the API, a # WARNING: partial capture header on the exported YAML, CaptureComplete: False on the CR. Export is never blocked; you just can't ship a partial profile by accident.
Export a CR, not a file
Part 1's network flow ends at YAML in a directory. Seccomp needs one more step, because the profile has to be a file on every node before a pod can reference it. That's the awkward bit most tutorials skip.
So the export is a SeccompProfile CR. Review it, commit it, apply it, and the controller already running on each node writes the file it describes.
That last part needs turning on, and it's off by default:
seccomp:
distribute: true
kubeletRoot: /var/lib/kubelet # k3s: /var/lib/rancher/k3s/agent/kubelet
Without it the CR applies cleanly and the file never appears on a node, because referencing a profile is a workload-availability decision the chart won't make for you. Get kubeletRoot right too; the default is wrong on k3s.
# observed syscalls: 55 (x86_64)
# capture: full, complete (1 contributing pod)
apiVersion: kguardian.dev/v1alpha1
kind: SeccompProfile
metadata:
name: deployment-checkout-api
namespace: seccomp-demo
annotations:
kguardian.dev/capture-complete: "true"
kguardian.dev/capture-level: full
spec:
defaultAction: SCMP_ACT_LOG # audit first, always
architectures: [SCMP_ARCH_X86_64]
syscalls:
- names: [accept, access, alarm, bind, ...]
action: SCMP_ACT_ALLOW
workloadRef: { kind: Deployment, name: checkout-api }

The export modal rendering the manifest, with Copy and Download
Every export starts at SCMP_ACT_LOG, for the reason at the top of this post. The file is kguardian/<ns>/<name>.json, rewritten in place, so version history lives in git where the CR does. The app team wires it up once and never touches it again:
securityContext:
seccompProfile:
type: Localhost
localhostProfile: kguardian/seccomp-demo/deployment-checkout-api.json

After applying: the CR is named, nodes read 3/3 Ready, drift is in sync
3/3 Ready counts the controllers that have written the file. Until it covers
the fleet, a pod can land on a node where the profile doesn't exist yet and
fail with CreateContainerError.
As in part 1, kguardian applies nothing. That step is yours.
Drift, and denials
Drift detail: five observed syscalls missing from the CR, with an Export updated CR button
That is real drift, not a staged one. Nobody touched the workload. It served
requests until it hit a path the first capture window had missed, and five
syscalls appeared that the committed CR knew nothing about: epoll_ctl,
futex, getpeername, getsockopt, setns. Observed went from 55 to 60. Had
that profile been enforcing, those five would have returned EPERM instead of
a warning in a table. That is what "observe for a full business cycle" means,
and how you find out you didn't.
Drift is the gap between what kguardian still observes and what your CR allows. drift.missing is what the workload calls and the CR doesn't allow, which is what breaks on promotion. drift.extra is what the CR allows but nothing ever called, which is where you narrow. The observed set is monotonic, so a syscall the workload stops making stays observed. A shrinking union would quietly narrow your profile toward breaking.
Denials are a different question. SCMP_ACT_LOG writes every would-be denial to the node's kernel audit log, which kguardian never used to read back, so confirming a profile was safe meant grepping dmesg on every node. That loop is now closed with a kprobe on audit_seccomp(), the function every loggable verdict passes through.
Drift |
Denials | |
|---|---|---|
| Source | eBPF observation | the kernel's own verdict |
| Fires when | workload makes a syscall the CR doesn't allow | a loaded filter denies or logs one |
| True with nothing referencing the profile | yes | no |
| Answers | is my profile falling behind? | is the kernel doing what I shipped? |
They come apart in both directions. The workload above drifted with zero denials, because nothing referenced its profile yet. And a profile with no drift at all can start denying the moment you promote it, because enforcement depends on timing, process ancestry and architecture in ways a set of syscall names can't capture.
Promotion, and how "clean" lies
The gate is DenialsObserved: False and status.denials.observed: 0. Then you change one line in git:
spec:
defaultAction: SCMP_ACT_ERRNO # was SCMP_ACT_LOG
Apply it, and restart the workload so the new filter loads. Two things have to
already be true or this does nothing at all: the pod template references the
profile, and Drift reads in sync. Promoting a drifted CR enforces an
allow-list you already know is short.
Next
Part 3, compute: the CPU and memory gauges, and the contention edge that says this pod starved that one. Then the eBPF half: probe design, network capture, the de-duplication that makes full cost what low costs, and a run-queue latency histogram at roughly 3% of a core.
kguardian is on GitHub. Docs at docs.kguardian.dev.
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes โ full credit and traffic to the original publisher.