The serving layer
Batching, prefill, prefix cache, speculative decoding
What is your GPU utilization during an ordinary business hour, not at peak?
Batching sized for the install, not the traffic
Requests queue while the card sits half idle
Prefill blocking decode
Tail latency spikes whenever a long prompt lands
No prefix caching on a shared preamble
Identical leading tokens recomputed on every request
Speculative decoding left on the table
Decode-bound workload with compute to spare