Multi-Tenant AI: Data Isolation
One bug in a multi-tenant AI application is categorically worse than the rest: answering customer A with customer B's data.
It's a breach, and it's quiet. The person who gets the wrong data may not realize what they're looking at. By the time anyone notices, the logs may be the only evidence. It's also easy to build by accident, since the default behavior of nearly every component in an AI stack is to return the best match, regardless of permission.
The principle
Identity travels with the request, all the way down, and no layer is permitted to default to everything. Every retrieval, every tool call, every cache lookup is scoped to the requesting tenant before it runs, not filtered afterward.
Filtering afterward is the specific mistake worth naming. Post-filtering leaks through ranking, through result counts, and through anything that reveals the shape of what was excluded.
Leak path one: retrieval
The most common failure. A vector store is queried for the top matches across the entire index, and the application filters by tenant before display.
The correct shape is a filtered search, where the tenant predicate goes into the query so the index only ever considers permitted vectors. Pinecone namespaces, Weaviate multi-tenancy, and Postgres row-level security with pgvector all give you this. Use it, and remove the possibility of forgetting:
// The raw client lives in internal/. A lint rule blocks imports of it from anywhere else.import { vectorStore } from "./internal/client"; export class TenantIndex { private constructor(private readonly tenantId: TenantId) {} // The only way to obtain an index is to present an authenticated context. static forRequest(ctx: AuthenticatedContext): TenantIndex { return new TenantIndex(ctx.tenantId); } async search(query: string, topK = 8): Promise<Match[]> { return vectorStore.query({ vector: await embed(query), topK, filter: { tenantId: { $eq: this.tenantId } }, }); }}What enforces the filter is that the unscoped client is unreachable: it lives under internal/, and a lint rule or package boundary blocks imports from there. A rule a developer has to remember is a rule that eventually gets forgotten, so the fix is to remove the option.
Leak path two: caching
A semantic cache keyed on query text alone is a cross-tenant leak. Customer A asks a question, the answer is cached, customer B asks something similar, and the cache serves A's answer.
The cache key must include the tenant, and usually anything else that scopes the answer, like role or region. If that makes the hit rate look disappointing, that's the measured hit rate.
More on cache economics in LLM Cost Control Engineering.
Leak path three: tools
An agent holding broad service credentials can reach whatever those credentials can reach, and eventually it will try. Propagate the end user's identity into every tool call and authorize per request at the resource, against that identity. The agent should see exactly what the user would see by logging in themselves, and nothing more.
The same argument at the protocol level: MCP in Production.
Leak path four: memory
Agent memory that accumulates across sessions needs the same scoping as retrieval. Both the episodic layer and any durable facts layer are per tenant, always. A shared memory store across tenants is a leak waiting for the right prompt.
How those layers separate: Agent Memory Architecture.
Leak path five: evaluation and training data
Eval sets and fine-tuning corpora assembled from production traffic mix tenants by default. If that data ever influences a model serving all tenants, a per-request isolation problem has moved into the weights, where you can't undo it. Treat any dataset built from customer traffic as requiring explicit consent and explicit scoping.
Leak path six: traces and logs
Prompts, retrieved documents, and tool results routinely land in a shared observability backend where access is granted by team without tenant isolation. That is a cross-tenant data path that never appears on an architecture diagram.
Redact or hash tenant payloads before they leave the request. Keep raw payloads in a store with the same access controls as the source data. Treat your tracing backend as a system that holds customer data, because it does.
Test for it adversarially
Add a class to your eval set that exists purely to break isolation: prompts that explicitly request another organization's data, that reference known identifiers from a different tenant, that try to get the agent to enumerate what it can see. The correct result is always a refusal or an empty result. If one of these ever passes, that's a stop-the-line failure.
Where these fit in the eval architecture: Evals for Production Agents.
Defense in depth
Even with every path above closed, add a final check on the way out: scan the response for identifiers belonging to tenants other than the requester. It should never fire. The day it does, you have found a path you didn't know existed, before your customer did.
Working on something like this?
Start a Conversation