riverfront ai labs
Agent reliability audit

Whether your agent survives production.

Most agent projects do not fail on model capability. They fail because nobody can say what correct means, the loop costs several times what anyone budgeted, and there is no defined behaviour for the day it is wrong.

$3,000 – $8,0005 – 10 days50% upfront


What we look at

Four layers, worked in this order. The first one decides whether the other three can be measured at all.

01

The evaluation set

Labelled real cases, and who signs off what correct means

If the agent answered badly this morning, how would you know by this afternoon?

  • No labelled set drawn from real cases

    Quality discussed in anecdotes and screenshots

  • Measured on the happy path only

    High scores, unhappy users

  • Correct defined by the builder

    No domain expert ever signed off the answers

  • Evals that do not run on change

    A suite that was run once, at launch

02

The loop

Context growth, retries, tool surface, task isolation

How many tool calls does a typical successful run take, and do you know the number or are you guessing?

  • Context growing without bound across a run

    Later turns slower, dearer and worse than early ones

  • Retries that repeat the failure

    The same tool call three times, then a give-up

  • A tool surface too large to choose from

    Wrong tool selected, or the right one used wrongly

  • No task isolation

    One long transcript doing four unrelated jobs

03

Unit economics

Cost per finished job, caching, routing, wasted runs

What does one completed job cost you, including the runs that failed on the way to it?

  • Cost measured per call, not per finished job

    A dashboard of token spend nobody can act on

  • Cache misses on a stable preamble

    A long fixed system prompt re-billed every turn

  • One frontier model for every step

    The same model classifying and reasoning

  • Failed runs billed and uncounted

    Spend that reconciles to no completed work

04

The failure path

Thresholds, escalation, permissions, decision logs

What is the agent allowed to do without a person, and who decided that?

  • No confidence threshold and no escalation

    Equal certainty in the answer and the guess

  • Irreversible actions without a gate

    Send, refund and delete on the same footing as read

  • No decision log that survives an audit

    Traces kept for debugging, retained for a week

  • Nobody paged when quality drifts

    Regressions found by a customer


What we report

Task success
The share of attempts that finished the job correctly, judged against the labelled set rather than against whether the run completed without erroring.
Cost / task
Total spend divided by successfully completed jobs, with failed and abandoned runs left in the numerator where they belong.
Turns / task
Tool calls and model calls consumed by a typical successful run. The number that tells you whether the loop is working or thrashing.
Escalation rate
How often the agent hands to a human, and how often it should have. The gap between those two is the finding.
p95 wall clock
End-to-end time for a whole task, not latency per call. What the person waiting on the result actually experiences.
Eval coverage
The proportion of production behaviour the suite actually exercises, and whether it runs on every prompt, tool and model change.

When not to call us

  • The agent is not in production yet. There is nothing to measure until real traffic has hit it, and a pilot with a friendly audience will tell you very little.
  • You want the agent built rather than examined. That is a different engagement with a different shape, and our consultancy practice runs it.
  • The problem is a model that cannot do the task at all. An audit finds recoverable reliability and cost; it cannot find capability that is not there.
  • Volume is low enough that a person doing the job costs less than the engagement. We will say so on the first call.

Send us a run that went wrong.

riverfront.aic@gmail.com

One trace, and what should have happened instead. If the fix is something you can do yourself in an afternoon, we will tell you what it is on the first call rather than sell you an audit.