Policy as Code for AI Workloads
AI workloads broke an assumption most clusters were built on: that what runs inside them is known in advance. A model server pulls weights from somewhere. An agent makes outbound calls to somewhere. A framework resolves packages at build time from somewhere. That is a supply chain and an egress surface that traditional workloads mostly didn't have.
Admission control is where you get to have an opinion about all of it before anything runs.
The mechanics, briefly
Kubernetes admission webhooks intercept a resource before it's persisted. Validating webhooks say yes or no. Mutating webhooks change the object on the way through. Policy engines like Kyverno and OPA Gatekeeper let you express those rules declaratively rather than writing webhook code.
Webhook or built-in. Kubernetes now ships ValidatingAdmissionPolicy, which evaluates CEL expressions in-process with no webhook at all. That removes a network hop and an entire failure mode: no webhook means no cluster-wide outage when the policy service is unhealthy. Use it for structural checks, which turns out to be a large share of what you need. Reach for a policy engine when you want image signature verification, external data lookups, mutation, or a policy library you would rather not rewrite in CEL.
The interesting question was never how to write a policy. It's which policies justify the friction.
The controls worth enforcing on AI workloads
Signed images only. Verify signatures at admission. Unsigned images don't run. This is the highest-value supply-chain rule in the set because it closes the "someone ran an image we didn't build" hole.
No mutable tags. No latest, no floating version tags in production. Digests or nothing, because a signature proves who built an image and does nothing to stop a tag being moved to a different signed one. A model server that changes under you without a deploy is an outage you can't reproduce.
Model provenance. If you mount weights, the source should be an approved registry or bucket, declared explicitly. A pod that can pull weights from an arbitrary URL is a pod that can run an arbitrary model. For example, allow weight volumes only from a named bucket prefix or the internal registry, and reject any pod whose init container or volume source points anywhere else.
Egress allowlists. The one teams skip, and the most important rule for agents. Default-allow egress means an agent with a code execution tool has a route to anywhere on the internet. Deny by default at the namespace level, allowlist the destinations each workload needs, and route through an egress gateway so there's a log.
GPU resource limits. Unbounded GPU requests are how one team's experiment starves production inference. Require limits and enforce them.
No model credentials in environment variables. Provider API keys belong in a secret store behind a short-lived token, and never in a pod spec that ends up in Git history.
A concrete policy
apiVersion: kyverno.io/v1kind: ClusterPolicymetadata: name: require-digest-and-approved-registryspec: rules: - name: images-must-use-digest match: any: - resources: kinds: [Pod] namespaces: ["inference", "agents"] validate: failureAction: Enforce message: >- Images must be pinned by digest and pulled from the internal registry. Replace the tag with an @sha256: digest. pattern: spec: containers: - image: "registry.internal.example.com/*@sha256:*" =(initContainers): - image: "registry.internal.example.com/*@sha256:*"Short, readable, and it removes an entire class of "what changed" incidents.
The failure mode to design against
Policies that block deploys nobody can debug. An engineer pushes, admission rejects with a message they don't understand, and the policy becomes an obstacle engineers route around.
Two habits prevent that. Run every new policy in audit mode first and look at what it would have blocked, because it's always more than you expect. And write the failure message as instructions rather than a rule name. Notice the message in the policy above tells you what to do. "Policy violation: require-digest-and-approved-registry" doesn't.
Policies break like code, so test them like code
Keep a set of manifests that should pass and a set that should fail, and run both in CI on every policy change. Discovering that your egress rule was too strict during an incident is an expensive way to learn it.
Working on something like this?
Start a Conversation