inference optimization
Batching policy, KV cache layout and reuse, speculative decoding, quantization — worked end to end until tokens per second per GPU stops being the constraint.
An engineering consultancy for teams running large models and agents in production. We profile what you have and report what we find, with the evidence attached.
Batching policy, KV cache layout and reuse, speculative decoding, quantization — worked end to end until tokens per second per GPU stops being the constraint.
When the framework has given everything it has, the time left is in the kernels. Warp-level profiling, fused attention paths, Triton, and PTX where the compiler will not cooperate.
Sharding that matches the hardware, autoscaling that understands token load rather than CPU, and the observability to prove the numbers hold under real traffic.
The evaluation set that decides whether it works, what one completed task actually costs once failed runs are counted, and what the system does on the day it is wrong.
Two tools that came out of the consultancy. Both run on hardware you already own.
for developersYour agent forgets everything at midnight.
Cameras you only look at after it already went wrong.
Fixed price, fixed scope, fixed finish line. Both are the usual way in.
For a serving stack that costs more than it should.
$2k – $5k5 – 10 days50% upfront
For an agent already in production, and what each completed job costs.
$3k – $8k5 – 10 days50% upfront
The same four stages whichever one you buy. Ten working days at the outside.
Whether there is enough headroom to be worth an audit, including when the answer is no. No fee.
We reproduce your workload and profile it end to end. Nothing changes in production.
Findings ranked by what they cost you, evidence attached. The ones we were wrong about get reported too.
The report, the cost model, and the changes in the order we would make them.
If your stack is not here, say so — the profiling work transfers.
For a serving stack: the model, the GPUs, your traffic shape, and the number that is bothering you. For an agent: one run that went wrong and what should have happened instead. You will get a reply from the person who would do the work.