Skip to content
Back to Insights
AI EngineeringBy KE Engineering Team

Versioning Prompts Like Code

Versioning Prompts Like CodeAI Engineering cover for Versioning Prompts Like CodebeforeafterAI ENGINEERINGVersioning PromptsLike Code// DIFF THE BEHAVIOR

A system prompt is program logic that happens to be written in English. It determines behavior, it has bugs, and a small edit can change outputs across every request. Yet in most organizations prompts live in a config file, a database row, or an admin UI, edited by whoever has access, with no review, no version, and no way to roll back.

That is how production code was managed in 1995. Here is what it looks like to catch up.

Prompts get semantic versions

Not every prompt change is the same size. Borrow the discipline from API versioning.

Patch changes fix wording without changing intended behavior. A typo. A clearer sentence that produces the same outputs on the eval set.

Minor changes add capability without breaking existing behavior. A new instruction for a case the prompt didn't previously handle, where all existing evals still pass.

Major changes alter behavior for existing inputs. Reordering instructions, changing the output format, adjusting tone in a way that shifts results.

The version number is a promise about what an eval run will show. A patch should produce a clean diff on the eval set. A major should not, and that is expected.

Prompts live in the repo

Not in a database. Not in a dashboard. In version control, next to the code that uses them, reviewed in the same pull request.

prompts/  support-agent/    system.md            # the prompt    CHANGELOG.md         # what changed and why, per version    evals/      cases.jsonl        # representative inputs      baseline.json      # expected outputs at the current version

A prompt change now goes through the same path as a code change: branch, edit, run evals, review the diff, merge, deploy. Someone reads it before it ships.

Review the eval diff before the text diff

A text diff of a prompt tells you what words changed. It tells you nothing about what outputs changed. The review artifact that matters is the eval diff: for each case in the set, the output before and after, with divergence flagged.

typescript
// CI step: run evals on the PR's prompt against the baselineconst results = await runEvals("support-agent", { prompt: prBranch });const diff = compare(results, baseline); if (diff.divergent.length > threshold[versionBump]) {  fail(`${diff.divergent.length} cases changed. Bump is ${versionBump}; expected ≤ ${threshold[versionBump]}.`);}

A patch that diverges on twenty cases isn't a patch. The tooling should say so before a human has to. A common version of this: a "patch" that changes "respond concisely" to "respond briefly and clearly," and the eval diff shows a fifth of the cases dropping a required field because the model read "briefly" as permission to omit it. Ephemeral Environments for AI Development is where that check runs against the real behavior surface.

Pin prompts to releases

A deployed system pins its prompt version the same way it pins its container digest. The running prompt is a known artifact, not "whatever is in the table right now." That makes "what is running" answerable and rollback a repeatable operation.

The model version is part of the contract

A prompt is tuned against a model. Change the model and the prompt's behavior changes even if its text didn't. Record the model snapshot against which the prompt was validated, and re-run evals when either side moves. A prompt version without a model version is half a version.

Roll a prompt out the way you roll out code

A prompt change is a production change, so it gets progressive delivery. Split traffic by session, never by request, so a single conversation doesn't flip versions mid-turn. Run online evals against both arms and gate promotion on them: structured output validity, tool-selection agreement, refusal rate. A canary that only watches p99 latency will happily promote a prompt that started refusing half of legitimate requests.

yaml
rollout:  artifact: prompt/support-agent@v15  baseline: prompt/support-agent@v14  strategy: canary  split_key: session_id          # never request_id  steps:    - weight: 5      hold: 30m      gates: [output_validity, tool_agreement, refusal_rate]    - weight: 25      hold: 2h      gates: [output_validity, tool_agreement, refusal_rate, approval_rate]    - weight: 100  rollback_on:    output_validity: { drop_pct: 2 }    tool_agreement: { drop_pct: 5 }    refusal_rate: { rise_pct: 3 }

Hold durations run longer than a code canary because behavior signals are noisier. Both arms are full release manifests: model, prompt, tool catalog, index build. Rollback swaps the manifest, because reverting the prompt while leaving the model in place doesn't restore the previous behavior. Tool-selection agreement is the earliest cheap signal: for the same input, do both arms pick the same tool with the same arguments?

The organizational change is the hard part

None of this requires a new vendor. It requires deciding that the people who edit prompts are doing engineering, and that engineering has a process. The pushback is usually that prompts need to be iterated quickly. They do, on a branch, with evals. The thing that needs to stop is iterating quickly in production.