Riverfront AI Labs

GPU performance engineering: inference, kernels, infrastructure

We make your inference stack faster.

Riverfront AI is a boutique engineering consultancy for teams running large models in production. We profile your serving stack, find where the GPU time and the money actually go, and fix it: batching, KV cache, quantization, kernels, topology.

What you get back is a lower cost per million tokens and a latency number you can put in front of a customer.

What we do

Three layers of the same problem. Most engagements touch all three.

LLM inference optimization

Most serving stacks leave a large share of the GPU idle and pay for it by the hour. We work the serving layer end to end, tuning batching policy, KV cache layout and reuse, speculative decoding and quantization until tokens per second per GPU stops being the thing holding you back.

Where that landsvLLM · TensorRT-LLM · SGLang · continuous & chunked-prefill batching · paged KV cache · prefix caching · speculative decoding · FP8 / INT8 / AWQ

GPU performance engineering

When the framework has given everything it has, the remaining time is in the kernels. We profile at the warp level and write what is missing: fused attention paths, custom Triton kernels, and PTX in the places where the compiler will not cooperate.

Where that landsCUDA · Triton · PTX · Nsight Compute · kernel fusion · memory coalescing · occupancy & register pressure · roofline analysis

AI infrastructure

Throughput on a benchmark rig is not throughput in production. We build the serving topology around the model: sharding that matches the hardware, autoscaling that understands token load rather than CPU, and the observability to prove the numbers hold under real traffic.

Where that landsdistributed inference · tensor & pipeline parallelism · Kubernetes · AWS · model serving · load-aware autoscaling · observability

Also from Riverfront

Riverfront AI Consultancy

Not every problem is a performance problem. Our consultancy practice builds the AI workflows themselves: document intake, back-office automation, retrieval and agents over your own data. Same standard of evidence, aimed at what the manual work costs you rather than what the GPUs cost.

Where that landsworkflow automation · agents & retrieval · data and ML engineering · strategy & fractional leadership

See the consultancy practice

Inference performance audit

Fixed price, fixed scope, fixed finish line. The usual way in.

Price
$2,000 – $5,000
Duration
5 – 10 days
Terms
50% upfront
  • Profile of your existing inference stack under production-shaped load
  • Bottleneck report: where the time goes, layer by layer, with the traces behind each claim
  • Batching configuration tuned to the shape of your actual traffic
  • KV cache sizing, layout and reuse review
  • Quantization recommendation with the accuracy trade-off stated plainly
  • Cost analysis: your current cost per million tokens, and the reachable one
  • Written final report your team can act on without us
  • Implementation support while the changes land

What we look for, and how we measure it

How the audit runs

Ten working days at the outside. You see findings as we get them, not at the end.

Day

  1. 0

    Scoping callNo fee

    Your model, your hardware, your traffic shape. We tell you whether there is enough headroom in the stack to be worth an audit, including the times when the answer is no.

  2. 1

    InstrumentDays 1–2

    We reproduce your workload under load and profile it end to end: request traces, GPU utilization, memory residency, kernel time. Nothing changes in production during this stage.

  3. 3

    DiagnoseDays 3–5

    Bottlenecks ranked by what they cost you, with the evidence attached to each one. You get measurements, not assertions. The ones we were wrong about get reported too.

  4. 6

    Report & land itDays 6–10

    The written report, the cost model, and the changes in the order we would make them. We stay on while your team lands the first ones.

What we work in

If your stack is not on this list, say so; the profiling work transfers.

Serving
vLLM · TensorRT-LLM · SGLang · Triton Inference Server · Ray Serve
Kernels
CUDA · PTX · Triton · CUTLASS · Nsight
Frameworks
PyTorch · JAX · ONNX Runtime
Infrastructure
Kubernetes · AWS · Docker · Terraform

When not to call us

Four situations where an audit is the wrong spend.

Contact

Tell us what your tokens cost.

riverfront.aic@gmail.com

Send the model, the GPUs you run it on, your rough traffic shape and the number that is bothering you: latency, throughput or the monthly bill. You will get a reply from the person who would do the work, usually within two business days.