Ephemeral Environments for AI Development
Preview environments per pull request are standard for web applications: push a branch, get a URL, review the change running. For AI features, most teams stopped at the code. The prompt change is reviewed as a text diff. The retrieval change is tested locally against a stale index. The model version is whatever the shared dev environment happens to point at.
That means the thing being reviewed isn't the thing that will ship.
What the environment has to include
An ephemeral environment for an AI change is the full behavior surface, spun up per branch:
- The application code, as usual
- The prompt and tool catalog from the branch, not from main
- The model pinned to the version the branch targets, with a scoped, short-lived credential
- A retrieval index built from a representative corpus, with the branch's chunking and embedding config
- The eval suite, wired to run against the environment and post results on the PR
The reviewer opens the preview URL and sees the actual behavior of the actual change. The eval results sit next to the code diff.
Keeping it cheap
Standing up a full index per PR sounds expensive. It is, if the index is the production corpus. It should not be.
Use a fixed, representative eval corpus: a few thousand documents chosen to cover the cases the evals exercise. It builds in minutes, it's repeatable, and every PR tests against the same corpus, so results are comparable. Production-scale retrieval behavior is validated in the canary stage, not the preview stage.
Model calls are the other cost. Cap them per environment and let the environment expire on a short timer. Route the preview through a cheaper tier only when the change is UI or plumbing, never when the prompt or model changed.
preview: ttl: 4h model: snapshot: from-branch # what the branch pins budget_usd: 5 # hard cap per environment index: corpus: eval-corpus@v3 # fixed, small, representative rebuild: on-config-change # not every push evals: suite: from-branch run_on: [push, manual] post_to: pull_requestThe review changes
With a working preview, prompt review stops being "does this wording look reasonable" and becomes "here are the eval diffs and here is the environment, go try it." Reviewers catch behavior changes the text diff hides. Product people can review AI changes without reading prompts at all.
Secrets and identity in previews
Preview environments are where credentials get forgotten. Each one gets a workload identity and short-lived tokens scoped to the preview, never a shared long-lived dev key. When the environment expires, its authority expires with it.
How you know it's working
You'll know it's working the first time someone rejects a PR because the preview behaved wrong, even though the text diff looked fine. A common version of this: a one-line prompt edit that reads as a clarification, and in the preview the agent starts answering a question the user didn't ask. That is a bug caught before production that no amount of prompt review would have found.
Working on something like this?
Start a Conversation