Skip to content
Back to Insights
AI EngineeringBy KE Engineering Team

Secure by Design: The Next Chapter of the Secure SDLC

Secure by Design: The Next Chapter of the Secure SDLCAI Engineering cover for Secure by Design: The Next Chapter of the Secure SDLC123456AI ENGINEERINGSecure by Design: TheNext Chapter of theSecure SDLC// FOUR PHASES. EVERY PROJECT.

Your system's behavior used to be a property of your code. Read the source, run the tests, and you knew what it would do. Put a model, a prompt, retrieved data, or an agent in it and that stops being true. A provider swaps a model snapshot. Someone edits a tool description in a config file. A document lands in the index. The code is byte-identical and the system does something new. Review before merge was built for the old case.

So we changed how we build. Every project runs the same four phases, Discover, Architect, Build, Enable, and none of them asks your team to work differently.

Discover: define the problem with the people who own it

We start with the people who own the problem: the workflows, the users, the decisions, the data, and the boundaries. For anything with a model or an agent in it, we also write down what good looks like for each interaction. Those expectations become its evaluations, and the boundaries become the security requirements the build is checked against.

Architect: challenge the design, then decide the safeguards

An independent reviewer questions the design before we build it: the assumptions, the failure modes, the security gaps, and what each model and agent is allowed to touch. Its job is to find what breaks, and the team defends the decisions that matter. A named person who isn't building it signs off. Then, still before any code, we decide the safeguards: the threat model, each protection the system needs with a drafted test, and an evaluation suite for every model interaction. During the build, tests get filled in, not invented.

Build: every change checked and evaluated

Every pull request runs tests, scans, a check for leaked secrets, a dependency check, a review from a model that didn't write the code, and the evaluations for any component it touches. You'll only see this in your normal code review. A common version of the failure this catches: an agent passes code review, then a prompt change a month later lets it call a tool nobody approved. So a change to a prompt, a model, a tool description, or retrieval logic takes the same path as a change to code, because it changes behavior the same way.

Evaluation is its own step. Every model interaction is tested against what Discover said good looks like: correctness, safety boundaries, grounding, tool selection, structured output, and cost where it matters. Failures from production go back into the suite, so it gets stronger as the system runs.

Before release, a different reviewer confirms the build is what was approved, every safeguard has its passing test, every evaluation meets its threshold, and each agent runs with exactly the freedom it was granted. The same named person signs the release.

Enable: run it, and own it

We put it into production and set up what it takes to run it: monitoring, logging, runbooks, documentation, and the evaluation process. We work alongside the team that will own it, on how it was designed, how to troubleshoot it, and how to change it safely. At handoff your people become the approvers, and they keep the patterns, controls, and operating model for the next system.

Three rules, all the way through

  1. No AI checks its own work. The AI that helps us build runs on one vendor and every reviewer runs on another, so a blind spot in one is less likely to be a blind spot in the other.
  2. Agents earn trust. Every agent we ship starts able only to suggest, gains freedom only through a decision someone signed, and anyone can take that freedom back.
  3. Every model interaction has its own evaluation. It runs on every change, and production failures feed it.

One file, if you want it

Everything above leaves a record: what we built, what was found and what we did about it, the evaluation results, who approved what and when, and the level each agent runs at. If anyone on your side wants to see it, it's one file. It describes the process we followed, not a certification of your system. Most projects never need it. The ones that do find the questions already answered.

Before there is code to secure

This is how we build whether or not anyone asks. The decisions that matter most, what a system may touch and what it may do, get made before there's any code to secure.