riverfront ai labs
wareywarey · a riverfront product

Your cameras already saw it.

Most warehouse cameras are a recording somebody watches after the fact. warey watches them live, on a box on your own network, and tells you while it still matters — seconds to an alert, everything else searchable in plain English, and no frame ever leaves the building.

in active buildrunning on live cameraseight cameras per box

343ms

frame to alert, p95

You know before the forklift has left the aisle. The budget was one second.

8

cameras on one box

Measured against the worst case H.264 allows, not against a friendly clip.

0

frames that leave your network

No cloud account, no upload, no monthly fee per camera.


The camera is not the problem

You already have the footage. What you do not have is somebody watching it at three on a Tuesday, or a way to find the ninety seconds that matter out of a fortnight.


What you get

One machine on your network, reading the cameras you already have. Two halves: one that has to answer now, and one that can take its time.

it tells you in seconds

Someone on foot in the forklift aisle. A blocked fire exit. A dock standing empty through a shift. Fixed rules rather than a model guessing, so the same footage always gives you the same answer.

you can ask it questions

Type “forklift moving cargo near the dock” and you get the moments back with their footage. Everything not worth an alert is still clipped, described and indexed — nobody presses anything.

it stays on your hardware

One box on your network, reading the cameras you already have over RTSP. Video is decoded on the GPU and written to your own disk. There is no cloud account and nothing to upload.

a restart does not cost you a day

Patch the OS mid-clip, lose power, reboot the machine. The unfinished work comes back on its own and gets done. Nothing waits for a person to notice it was dropped.


What it catches

Twelve rule kinds ship. These are the ones sites ask for first.

  • person on foot in a vehicle aisle
  • blocked fire exit
  • dock idle through a shift
  • after-hours movement
  • zone occupancy, by the minute
  • loitering at a bay

How a camera is set up

No zone is drawn in advance and no rule ships pointed at your floor. A camera warey has not been set up against runs with no zones and says so.

  1. step 1

    you answer four questions

    What kind of place this is, when it is worked, what is dangerous here, and what you want to be told about. A system that guesses those watches the wrong things.

  2. step 2

    it watches the camera

    A few minutes at two frames a second, recording where people and vehicles actually go, and naming the areas it can see — the dock, the walkway, the racking.

  3. step 3

    the rules arm themselves

    A dangerous-by-default set arms for the areas genuinely found, then your own sentences become rules. Anything that cannot be expressed honestly is refused with a reason.

  4. step 4

    nothing is saved until you say yes

    Every proposed area is drawn over a frame your own camera produced, in a browser, with its evidence beside it. Turn one off and its rules go with it.


What it refuses to do

A monitoring system earns its place by being believed. The fastest way to lose that is a confident alert about something it could not actually see.

it will not guess a distance

Forklift proximity needs a calibrated camera. On an uncalibrated one the rule is not approximated, and not even offered.

why
Turning pixels into metres requires calibration. Approximating it would put a permanently disabled rule on the dashboard, and that reads as coverage which does not exist.

it will not invent a capability

Asked to detect someone falling over, it says it cannot — that needs a pose model this build does not have.

why
Four of the twelve rule kinds refuse for reasons like this, and each refusal is shown on the dashboard rather than buried.

a number arrives with its coverage

Occupancy is rolled up per zone per minute, and every row records how many samples it saw.

why
A minute averaged over three looks is not the same evidence as one averaged over 240, and a camera that was down does not get to read as a quiet floor.

the same camera gives the same answer twice

Set the same camera up again and you get the same areas and the same rules, to the letter.

why
That sounds obvious and was not. The models that read your floor are sampled, so without a fixed seed one setup was a draw rather than an answer — and nobody could tell a change in the software from a change in the dice.

it will not quietly stop indexing

A clip it could not read is recorded as failed, with the reason, and counted separately from one whose retention had already expired.

why
A queue that silently drops work looks exactly like a quiet week.

What it measures

Every number here was taken on an 8 GB laptop GPU under load, with the database and the stream server running, and can be re-run on yours.

343ms

frame to alert, p95

Against a one-second budget.

how it was measured
Measured on a single 8 GB laptop GPU with the database and the stream server also running. Most of that is the wait for the next sample, not the model.

111fps

detector throughput

Three and a half times what eight cameras actually need.

how it was measured
Eight cameras at four samples a second require 32 frames per second. 111 is the figure under load, on the detector alone.

1485MiB

decode memory, eight streams

The tightest number in the system, and the one that sets the camera ceiling.

how it was measured
Measured against the worst case H.264 allows rather than against a friendly clip.

0.848

person detection, AP@0.5

An upper bound from public data, not a promise about your floor.

how it was measured
Against labelled ground truth. The number that matters is the one measured after your own cameras are up.

118ms

two cameras, sustained

Against a 250 ms budget with both live views open. With nobody watching, 102 ms.

how it was measured
Zero overruns across a full run on real fixed-camera footage.

21/21

alerts clipped and indexed

Two cameras over RTSP for a hundred seconds. Every alert got its footage.

how it was measured
21 alerts raised, 21 clips cut from the recording and attached to them, every one queued and indexed into something searchable. Zero missed loop deadlines, and zero frames the GPU could not decode.

225ms

eight cameras, every view open

Inside the 250 ms budget, but only just — a 1.1× margin, not a comfortable one.

how it was measured
The worst case this system has, measured three times rather than estimated once. It is the number that decides how many cameras one box takes.

Watching a floor is watching people

Which is why the limits are in the design rather than in a policy document.

no faces, no re-identification

A tracked person gets an anonymous number, local to one camera, that dies when they leave the frame.

detail
No face recognition, no gallery, and no matching a person across two cameras.

masked areas are masked in the pixels

Regions you nominate are blacked out before anything is stored or analysed, not filtered afterwards.

every look is logged

Every view, export and search is written to an access log with who did it and why.

footage expires on a schedule

Continuous video 14 days, event clips 90 days, and the searchable index two years.

detail
Retention runs ahead of the disk filling rather than waiting for it, because the recording that fails on a full disk is the one of the incident that filled it.

Tell us about your cameras and we will tell you what it costs.

Resolution, codec and count decide the hardware, and they are the first thing we would ask. If yours are 1080p H.264, the numbers above are already yours.

Start a conversation

The audits and the consultancy are a different door — riverfront ai labs if you wanted those instead.