riverfront ai labs
gpu performance · agent reliability · ai infrastructure

We fix the AI you already run.

An engineering consultancy for teams running large models and agents in production. We profile what you have and report what we find, with the evidence attached.


What we do

inference optimization

Batching policy, KV cache layout and reuse, speculative decoding, quantization — worked end to end until tokens per second per GPU stops being the constraint.

gpu performance

When the framework has given everything it has, the time left is in the kernels. Warp-level profiling, fused attention paths, Triton, and PTX where the compiler will not cooperate.

ai infrastructure

Sharding that matches the hardware, autoscaling that understands token load rather than CPU, and the observability to prove the numbers hold under real traffic.

agent reliability

The evaluation set that decides whether it works, what one completed task actually costs once failed runs are counted, and what the system does on the day it is wrong.


Two audits

Fixed price, fixed scope, fixed finish line. Both are the usual way in.

inference audit

For a serving stack that costs more than it should.

$2k – $5k5 – 10 days50% upfront

  • Your stack profiled under production-shaped load
  • Bottleneck report with the traces behind each claim
  • Cost per million tokens now, and the reachable one

What we look for

agent audit

For an agent already in production, and what each completed job costs.

$3k – $8k5 – 10 days50% upfront

  • A labelled eval set built from your real runs, failures included
  • Cost per successfully completed task, abandoned runs counted in
  • The eval harness running in your CI, owned by you

What we look for


How it runs

The same four stages whichever one you buy. Ten working days at the outside.

  1. day 0

    scope

    Whether there is enough headroom to be worth an audit, including when the answer is no. No fee.

  2. days 1–2

    instrument

    We reproduce your workload and profile it end to end. Nothing changes in production.

  3. days 3–5

    diagnose

    Findings ranked by what they cost you, evidence attached. The ones we were wrong about get reported too.

  4. days 6–10

    hand back

    The report, the cost model, and the changes in the order we would make them.


What we work in

If your stack is not here, say so — the profiling work transfers.

  • vLLM
  • TensorRT-LLM
  • SGLang
  • CUDA
  • Triton
  • PTX
  • Nsight
  • PyTorch
  • Claude Agent SDK
  • LangGraph
  • MCP
  • Langfuse
  • OpenTelemetry
  • Kubernetes
  • AWS
  • Terraform

When not to call us

  • You have not shipped yet. There is nothing to profile until real traffic exists.
  • You want somebody to own the system permanently. We measure, document, hand back, and leave.
  • The job is to build the thing rather than examine it — that is our consultancy practice, and it runs differently.
  • The spend is small enough that we would cost more than a year of the savings. We will say so on the first call.

Tell us what it costs you.

riverfront.aic@gmail.com

For a serving stack: the model, the GPUs, your traffic shape, and the number that is bothering you. For an agent: one run that went wrong and what should have happened instead. You will get a reply from the person who would do the work.