riverfront ai labs
Inference performance audit

Where the GPU time actually goes.

An audit is five to ten days of profiling your production inference stack and reporting, with evidence, what is costing you latency and money.

$2,000 – $5,0005 – 10 days50% upfront


What we look at

Four layers, worked in this order. Findings from the earlier ones usually make the later ones unnecessary.

01

The serving layer

Batching, prefill, prefix cache, speculative decoding

What is your GPU utilization during an ordinary business hour, not at peak?

  • Batching sized for the install, not the traffic

    Requests queue while the card sits half idle

  • Prefill blocking decode

    Tail latency spikes whenever a long prompt lands

  • No prefix caching on a shared preamble

    Identical leading tokens recomputed on every request

  • Speculative decoding left on the table

    Decode-bound workload with compute to spare

02

KV cache and memory

Context length, precision, fragmentation, paging

How much of your GPU memory holds live cache, and how much is reserved for a context length nobody ever sends?

  • Context length set far above real usage

    Large reserved cache the workload never fills

  • Cache held at full precision

    KV in FP16 while the weights are quantized

  • Fragmentation and reservation waste

    Allocator failures well below nominal capacity

  • Precision chosen once, at the start

    FP16 weights on hardware with FP8 units

03

Kernels

Attention backend, fusion, coalescing, occupancy

Do you know whether your decode step is compute-bound or memory-bound? If not, that is the first thing an audit tells you.

  • Wrong attention backend for the workload

    Kernel time dominated by attention at your sequence lengths

  • Unfused pointwise chains

    Memory traffic far above the arithmetic requires

  • Uncoalesced access in custom operators

    Achieved bandwidth well under the roofline

  • Occupancy capped by register pressure

    Few active warps per SM despite a large grid

04

Serving infrastructure

Autoscaling, cold starts, routing, parallelism

When traffic doubles, how long before you are serving it at your normal latency?

  • Autoscaling on CPU metrics

    Replicas added late, removed early, or both

  • Cold starts pulling weights over the network

    Minutes between a scale-up decision and a served request

  • No length-aware routing

    Long and short requests contending on the same replica

  • Parallelism mismatched to the hardware

    Tensor parallelism crossing a slow interconnect


What we report

TTFT
Time to first token, which is what a user perceives as responsiveness. Dominated by prefill and by queueing, which are two very different fixes.
TPOT
Time per output token once streaming starts. Memory-bandwidth bound in almost every case, and the number batching trades against.
Throughput
Output tokens per second per GPU. The number your bill is actually a function of.
Goodput
Requests per second served inside your latency target. Throughput without this is a number you cannot sell to anyone.
Cost / M tokens
Instance cost divided by delivered tokens. The translation layer between an engineering change and a finance conversation.
MFU
How much of the hardware's arithmetic capability you are actually using, against how much you are renting.

When not to call us

  • You have not shipped inference yet. There is nothing to profile until real traffic exists. Build first, and call us when the bill starts to hurt.
  • The problem is answer quality, not speed. That is an evaluation problem, and our agent reliability audit is the one that goes after it.
  • You want someone to own the stack permanently. We optimize, document, hand back, and leave.
  • Your workload is small enough that a larger instance costs less than the engagement. We will say so on the first call.

Send us the profile you already have.

riverfront.aic@gmail.com

Even a rough one: model, GPUs, concurrency, current latency. If we can see from that alone that there is little headroom left, we will tell you on the first call rather than sell you an audit.