Retail has never recorded more and shrink still went up. Retail loss prevention video only changes the number when it moves from documenting a loss to interrupting one — and that shift is narrower, and more honest, than most pitches admit.
Coverage went up. Resolution went up. Retention went up. And shrink went up with all of it — the National Retail Federation put it at 1.6% of sales for FY2022, $112.1 billion, against 1.4% the year before.
That is the uncomfortable starting point for any conversation about retail loss prevention video. A decade of investment in seeing more did not produce losing less, and the reason is structural rather than technical: a camera that records is an evidence device. Evidence is produced after a loss, and a loss that has already happened is not recoverable by knowing more about it.
Which means the useful question is narrow: what can video change while the event is still in progress?
What the Numbers Honestly Support
Before claiming anything, it's worth being straight about how soft this evidence base is — because the retail industry's own headline statistics have not held up well.
In December 2023 the NRF retracted its widely-quoted claim that organised retail crime accounted for nearly half of shrink. The figure had been assembled by attaching a $45 billion estimate from Senate testimony to the federation's own shrink total, and it contradicted the federation's own survey, which attributed 37% of losses to external theft of every kind — not just organised groups. The NRF acknowledged the difficulty of assembling an accurate and agreed-upon dataset.
Take two things from that. First, treat every dramatic retail theft number — including any a vendor quotes at you — as contested until you see its derivation. Second, and more usefully: external theft is a minority of shrink. The rest is internal theft and process failure, and no camera in the aisle addresses a receiving error, a markdown that never got recorded, or a supplier discrepancy.
So the honest scope of real-time video is a slice of a minority. That is not an argument against it. It is an argument for sizing it correctly, and for being suspicious of anyone who presents cameras as a solution to shrink rather than to one component of it.
Within that slice, though, the timing question is everything.
The Window Is Narrower Than the Incident
Retail theft is not a single moment. It's a sequence — selection, concealment, movement toward the exit, departure — and the value of an alert collapses as that sequence progresses, because the range of safe available responses collapses with it.
| Moment | What the system can do | What it's worth |
|---|---|---|
| Selection — item goes in a bag | Raise an alert on a defined behaviour in a defined zone | Highest. Nothing has been lost yet |
| On the floor, pre-exit | Send an associate to offer service nearby | High, and safe — presence ends most attempts |
| At the door | Alert staff who are told not to intervene physically | Low. Policy usually forbids the only action left |
| After exit | Produce footage for a report or a case file | Evidentiary only. The loss already happened |
The table is the whole argument. Almost every deployment concentrates its detection at the door, which is the point where the merchandise is already concealed, the subject is committed, and store policy almost universally forbids staff from physically intervening. An alert delivered there is a notification that a loss is leaving the building.
Move the same detection earlier and the response changes from confrontation to service. An associate who walks into the aisle and asks whether someone needs help finding a size is not accusing anybody, is not at risk, and has just removed the conditions the attempt depended on. That intervention is safe, deniable, lawful, and it happens before anything has been lost.
Which is a latency requirement disguised as a staffing decision.
Why Real-Time Retail Loss Prevention Video Fails in Practice
An alert about the health-and-beauty aisle that reaches someone ninety seconds later describes a place the subject has left. The response loop in retail is the same closed loop as in any intervention — observe, decide, act, see the result — and it is subject to the same budget covered in the intervention latency piece. Retail just spends it differently.
The differences matter. There is no control room. The person who has to act is a floor associate holding a handheld, doing three other jobs, and the alert competes with a customer standing in front of them. Verification usually happens on a phone screen in poor lighting. And there is rarely a second chance, because the subject is walking.
So the practical target is not sub-second glass-to-glass for its own sake — it's that from event to an associate looking at the clip on a handheld, the total is short enough that the aisle reference is still true. If your alert path routes video to a cloud, runs inference, sends a push notification and then streams a segmented clip to a phone, measure the whole chain end to end before promising anyone an interruption.
And then measure the thing that actually decides whether any of it gets used.
The Alert Budget Nobody Sets
A monitoring centre has operators whose job is adjudicating alerts. A store does not. The person receiving retail alerts has a primary job that isn't security, which means the tolerable alert volume is dramatically lower than in a staffed operation — and the operator-to-camera ratio arithmetic lands harder here, because the denominator is somebody's spare attention rather than their salary.
Practically: an associate who receives six alerts in an hour and finds five of them worthless stops opening the seventh. Nobody logs that decision, no dashboard reports it, and the system continues generating alerts into a channel that has been quietly abandoned. The failure looks exactly like success right up until someone asks how many alerts led to an intervention.
Which sets a hard design constraint. Precision matters more than recall in retail, by a wide margin. A system that surfaces two high-confidence events a shift and gets acted on beats one that surfaces forty and gets muted, even though the second one detects more.
That constraint points at where to aim the cameras.
Where Real-Time Actually Pays
- Self-checkout. The highest-yield zone in most stores, because the loss mechanism is procedural rather than furtive — non-scans, mis-scans, produce substitution — and it correlates directly with a data stream you already own. Video paired with POS exception data gives an alert with a receipt attached, which is both more accurate and easier to act on politely.
- High-value fixtures, at selection. A defined zone, a defined dwell threshold, and a service response. Detect the moment of selection, not the moment of exit.
- Back-of-house and receiving. Unglamorous, and it addresses internal theft and process failure — the majority of shrink that aisle cameras never touch. The intervention here is process correction, not interception.
- Repeat-subject recognition across visits, carefully. Genuinely useful for organised activity, and the area with the heaviest legal and reputational exposure. Biometric identification is regulated very differently across jurisdictions, so this is a legal decision before it is a technical one.
Notice that two of the four are not about catching anybody. The biggest available wins in retail shrink are procedural, and a camera's role there is to show you what your process actually does rather than to interrupt a person.
Which changes what the infrastructure has to do.
What This Demands of the Video Layer
A chain is a multi-site problem before it is a camera problem. Hundreds of stores, each with its own uplink and its own bandwidth ceiling, all needing to reach central loss-prevention — and connecting many sites to one operations centre isn't a bandwidth problem, it's a topology one. Pushing every store's full-resolution stream to head office is the design that fails first.
The live leg has to reach a handheld quickly, which means browser-deliverable video without a plugin and without a multi-second segment buffer — the transport problem covered in why RTSP cameras still can't talk to browsers. And the archive leg still has to exist underneath it, because the case file, the insurance claim and the prosecution all draw on recording at scale and its compliance constraints.
Which puts the sourcing decision in familiar terms.
Build, Buy, or Deploy
- Build on open source. go2rtc or MediaMTX at each store for ingest and low-latency delivery, with your own detection and POS correlation on top. You own the per-store cost, which matters when the estate is hundreds of sites. Right when loss prevention is a core competency and you have engineering.
- Use a managed cloud relay. Fast to deploy across a chain, nothing per-store to run — but per-stream pricing multiplied by store count is the cost curve that ends these projects, and customer-facing footage in a third party's cloud is a privacy question you will be asked.
- Deploy a commercial platform. A full stack — per-site ingest, low-latency delivery to a handheld, recording and retention — on-premise or as a managed cloud deployment under a flat license. Samvyo is one such option: based on SFU architecture, with the media path, TURN and recording under your control, which matters when footage of your customers is the asset being stored. Where it doesn't fit: a handful of stores where a VMS and an occasional footage pull are genuinely enough.
Three routes, one test — can an associate see the aisle while the person is still standing in it?
The Bottom Line
Recording more didn't reduce shrink, and it was never going to: footage is produced after the loss it documents. Real-time retail loss prevention video is worth building, but size it honestly — external theft is a minority of shrink, the industry's own headline theft statistic was retracted in 2023, and process failure and internal loss are untouched by an aisle camera. Within its real scope, the whole game is moving detection from the door back to the moment of selection, because that is the only point where the available response is a safe one: an associate offering help, before anything has been concealed. Get the alert there while the aisle reference is still true, keep precision high enough that nobody mutes the channel, and put the biggest effort where the biggest losses actually are — which is usually self-checkout and the stockroom, not the shop floor.
What's Next?
The response loop this depends on, and how much delay it can absorb before an alert is just a recording with extra steps, is covered in Talk-Down at Ten Seconds Is Theater.
For whether anyone can actually adjudicate the alerts you generate, the operator-to-camera ratio does the arithmetic.
And for the distinction between recording and intervening in general, Proactive vs Reactive Surveillance draws the line.
Frequently Asked Questions
Does retail loss prevention video actually reduce shrink?
Recorded video on its own has a poor track record — coverage and retention rose across the industry while shrink rose to 1.6% of sales in FY2022. Video reduces loss when it triggers an intervention before concealment, not when it documents one afterwards. And it only ever addresses external theft, which the NRF's own survey puts at a minority of total shrink.
What percentage of retail shrink is actually theft?
Less than commonly claimed, and contested. The NRF attributed 37% of losses to external theft of all kinds, and in December 2023 retracted its separate claim that organised retail crime accounted for nearly half of shrink, acknowledging the difficulty of assembling accurate data. Internal theft and process failure make up the remainder.
When is the best moment to intervene in a retail theft?
At selection, before concealment. That's the only point where the available response is safe and non-confrontational — an associate offering assistance nearby. By the exit, the merchandise is concealed, the subject is committed, and most store policies forbid physical intervention, so an alert there is effectively a notification of a loss in progress.
Why do store staff stop responding to video alerts?
Because precision is too low and the recipient has another job. An associate who finds most alerts worthless stops opening them, and nothing in the system records that they've disengaged. In retail, precision matters far more than recall — two acted-on alerts a shift beats forty that get muted.
Is self-checkout worth monitoring with video?
It's usually the highest-yield zone in the store, because the loss mechanism is procedural — non-scans and mis-scans — and correlates directly with POS exception data you already collect. Pairing video with that data produces alerts with evidence attached, which are both more accurate and easier to handle politely.
What video infrastructure does real-time retail LP need?
Per-store ingest with browser-deliverable low-latency video reaching a handheld, plus retention for case files — across an estate of many sites without pushing full streams to head office. Samvyo is based on SFU architecture and keeps the media path, TURN and recording under your control, which matters when the footage is of your customers.