Skip to content
Back to Insights
PlatformBy KE Engineering Team

Serving Models in Your Own Cluster

Serving Models in Your Own ClusterPlatform cover for Serving Models in Your Own Clusterqueuegpu replicasPLATFORMServing Models in YourOwn Cluster// COLD START IS THE PROBLEM

At some point the hosted API bill, a data residency requirement, or a latency target pushes a team to run inference themselves. That is a reasonable decision. It's also a bigger one than it looks, because the moment you serve your own models you are running a GPU platform, and GPU platforms don't behave like the other workloads in your cluster.

Cold start drives the entire architecture

A normal service starts in seconds. An inference replica pulls a container image, then loads tens of gigabytes of weights into GPU memory. On a cold node with weights coming from object storage, you can wait several minutes before the first token.

Inference request pathGateway to queue to a pool of inference replicas, with a node-local weights volume feeding the replicas and object storage staged out of the critical path.INFERENCE PATHGATEWAYQUEUEwait = signalscale on queue waitREPLICA 1REPLICA 2REPLICA 3WEIGHTS VOLUMEnode-local, warmobject storage staged ahead of time// REPLICAS READ WARM WEIGHTS FROM LOCAL DISK
Fig. 1: Keep weights off the critical path.

That single fact determines most of what follows. The options, roughly in order of how much they help:

Keep weights off the critical path. Stage them on a node-local volume or a shared filesystem that's already warm. Pulling from object storage on every cold start is the slowest possible choice and the easiest to accidentally ship. For example, a replica that pulls tens of gigabytes from a bucket on every start spends minutes on the download alone; staged on a warm volume, that cost is paid once per node.

Separate the image from the weights. Baking multi-gigabyte weights into a container image makes the image pull the bottleneck and turns every model update into a full rebuild. Keep the runtime image small and mount the weights.

Keep a warm floor. Scale-to-zero looks excellent on a cost dashboard and terrible on a latency dashboard. If the workload is user-facing, hold a minimum replica count. Scale-to-zero belongs to batch jobs and internal tools, not to anything a person is waiting on.

GPU utilization is the wrong autoscaling signal

A replica can sit at high utilization serving one request badly, or at moderate utilization comfortably batching twenty. Scaling on it produces oscillation: you add capacity that doesn't help, then remove it while requests are still queuing.

Scale on queue wait instead. It's what users feel, and it maps cleanly to the decision you want to make.

yaml
apiVersion: keda.sh/v1alpha1kind: ScaledObjectmetadata:  name: inference-workersspec:  scaleTargetRef:    name: inference  minReplicaCount: 2          # warm floor, not zero  maxReplicaCount: 24  cooldownPeriod: 300         # scale down slowly, cold starts are expensive  triggers:    - type: prometheus      metadata:        serverAddress: http://prometheus.monitoring:9090        # queue wait, not GPU utilization        query: |          max_over_time(inference_queue_wait_seconds{quantile="0.95"}[2m])        threshold: "0.5"      # below your first-token target; queue wait is only part of it

Both the warm floor and the long cooldown are deliberate. Scaling down aggressively is how a cost optimization becomes a latency incident.

Batching is a latency decision wearing a throughput costume

Continuous batching is what makes self-hosting economically viable. It's also what makes your tail latency worse. The harder you batch, the better your tokens per second per GPU, and the longer an unlucky request waits to join a batch.

vLLM and SGLang both handle this well and they schedule differently, so their defaults produce different tail behavior on the same workload. Whichever you run, the defaults were tuned for someone else's traffic. Write down your first-token target as a number, then tune batch size and scheduling policy against it. A team that has not written that number down is optimizing for whichever metric someone happened to open first.

Multi-model is where the capacity math gets concrete

One model per node pool is simple and wasteful. Multiplexing several models onto shared GPUs is efficient and complicated.

The middle path most teams land on is one base model per pool with LoRA adapters swapped per tenant or per task. The expensive weights stay resident and variation becomes cheap. This is the right default for anything that looks like per-customer specialization.

When you need hard isolation between models on one physical GPU, MIG partitioning gives you that, with the tradeoff that partitions are fixed and can't be resized without draining the node. That rigidity is the cost of the isolation, and it's worth paying only when the isolation is a firm requirement.

If you need several large models live simultaneously, do the capacity math before you commit. Say three models each need one 80GB GPU at peak, and their peaks don't overlap. Time-sharing looks like it saves a GPU. But weights and KV cache share the same memory, and one coinciding peak means both models need their full allocation at once. Plan for one GPU set per model plus cache headroom. That is a fleet.

What to instrument

Time to first token at p50 and p99. Tokens per second per replica. Queue depth and queue wait. GPU memory in use versus allocated. KV cache utilization, which is what caps concurrent requests per replica.

Add KV cache eviction rate to that list. When the cache starts evicting under load, effective concurrency drops before any utilization metric moves. That is the specific reason teams watch latency degrade while every dashboard looks healthy.

When not to do this

Low or spiky volume, where you pay for idle GPUs between bursts. No platform team, because this becomes someone's permanent job. And any case where the hosted bill is still smaller than the engineering time it takes to run this well. Self-hosting starts to pay off when volume is steady and predictable enough that idle GPU time stays low and the platform work amortizes across many requests.