PerspectiveAugust 21, 2026·7 min read

Less Than 1% of the World's Video Is Ever Watched

Less Than 1% of the World's Video Is Ever Watched

The world has already deployed the sensors. It has not deployed the understanding.

Three numbers frame the category. They come from NVIDIA's published vision-AI materials, which is where most of the industry gets them: more than 1.5 billion enterprise cameras deployed worldwide, producing roughly 7 trillion hours of video per year, of which less than 1% is ever reviewed by a human. Video is already more than half of internet traffic, and under 1% of that is analyzed for insight.

Seven trillion hours a year is about 800 million years of continuous footage — recorded, held for a compliance window, and then overwritten unseen. That is our arithmetic on their installed-base figure, and it is the only way we've found to make the number mean anything.

The cameras were bought for good reasons — deterrence, insurance, compliance, forensics — and they keep being bought. Camera hardware prices fall around 5% a year structurally, and AI-capable cameras now ship at $300–800 with inference silicon on board. The constraint was never sensing. It was that turning footage into answers has historically required either a human watching, or per-camera analytics appliances so expensive and so narrow that they were reserved for a handful of "important" views. Both approaches leave essentially all recorded reality untouched.

The unwatched 99% is not noise#

It contains the answers to questions operators already ask, and today answer from memory, from paperwork, or not at all:

  • When did the pallet in bay 12 leave, and on whose truck?
  • Who propped the fire door between 2 and 4 a.m.?
  • How many near-misses did that forklift corner produce this month?
  • Was the wet-floor sign up before the fall that became a claim?
  • Which entrance did the person in the red jacket use, and was a vehicle waiting?

Every one of these is a retrieval question over events that were in frame for some camera. None of them requires a new sensor. What they require is that the footage be converted, at write time, into something queryable — and until recently the economics of that conversion were prohibitive. In our own production pipelines, embedding generation has run roughly 90% of total pipeline cost (first-party, and we label it that way wherever we use it). Decode, detect, store, and search together are the rounding error; the embedder is the meter that runs all day.

Three ways the industry compensates, and why each is breaking#

Nobody accepts the unwatched 99% happily. There are three standard workarounds, each a substitute for the same missing primitive, and each failing on its own terms.

Posted labor — watch a subset, live. The supply problem comes first. The U.S. Bureau of Labor Statistics projects approximately zero net growth in the 1.27-million-person guard workforce through 2034, with ~162,000 openings a year that exist purely to backfill turnover running at 100–200% annually — and a single 24/7 post costs $150,000–220,000 a year to staff at all. Then, even where you can fill the seat, attention doesn't scale: one person covers a wall of tiles, and the standing figure for how much of what reaches them is noise — 94–98% of police alarm calls are false — is a statement about the limits of human attention before it is a statement about anything else.

Sampling review — watch a subset, later. Loss prevention, safety, and QA teams spot-check recorded footage on a schedule. Sampling works when incidents are frequent relative to the sample rate, and fails silently for exactly the events that matter: rare, brief, and spread across cameras. The staffing to raise the sample rate doesn't exist — 65% of hotels report staffing shortages and 71% can't fill open roles, and surveillance review is precisely the discretionary shift that goes unfilled first.

Forensics — watch a subset, too late. When something serious happens, someone finally watches, linearly, at 1× or 4×, across every camera that might have seen it. Detectives describe manual review consuming tens of hours per case. The value is real — one published study found video evidence raised case solvability for firearm assaults by 442%, a larger effect than DNA or eyewitness testimony — but it is bought at a labor price that caps how often anyone can afford to pay it, and it is purely retrospective. Nothing about a forensic review prevents the next incident.

All three have the same shape: spend human hours converting footage into answers, after the fact of recording. The labor is the bottleneck, and the bottleneck is structural.

SectorThe labor constraintWhat it costs
Security~0% net workforce growth through 2034; 100–200% annual turnover$150–220K per 24/7 guard post; 94–98% of alarm calls false
Hospitality65% of hotels understaffed; 71% can't fill open roles4–6% of revenue lost to theft, largely unreviewed
MunicipalMulti-camera review measured in detective-days per caseSolvability left on the table (+442% when video is found)
WarehouseSafety observation is manual and sampled$84M/week in U.S. non-fatal injury costs; cargo theft +60% YoY
ConstructionSites are unmanned nights and weekends by definition$300M–1B/yr U.S. equipment theft; ~80% never recovered

This is not a wage problem that a recession fixes. It is demographic (the guard workforce is flat at any plausible wage), attentional (humans cannot watch tiles), and economic (the labor cost per watched camera-hour has no curve working in its favor). The cost of machine watching, meanwhile, is falling on four axes at once — which is a separate post.

The other sensors are already installed too#

Every facility in that table also runs badge readers, RFID scanners, GPS telematics, alarm panels, and transaction systems. That data is already structured and nearly free to ingest. And the questions operators actually need answered usually live in the join between a camera and one of those systems: tailgating, fictitious pickups, cloned credentials, equipment moving off-site while nobody appears on the yard cameras. No amount of camera-only intelligence finds those events, because the event doesn't exist in any single modality.

What "watched" would even mean at fleet scale#

The naive mental model — run a model on every frame and store what it sees — fails arithmetic before it fails anything else. A fleet analyzing at 5–15 fps with several detections per frame persists billions of records per camera-year, nearly all of them repeats: the same parked forklift, the same shelf, frame after frame. Modeled honestly, a naive every-detection-is-a-write policy overstates real storage volume by roughly 20×. (That ratio is a planning band from our own scale modeling, not a measured production figure.)

The workable definition is event-driven: you don't store frames, you store what happened. Detection and tracking run continuously, but a vector is written only at meaningful boundaries — a track entering, an appearance changing, a better view replacing a worse one, a periodic refresh. Under that policy a camera writes on the order of 0.1–10 vectors per second depending on scene density, and the queryable unit becomes the episode — something appeared, persisted, moved, left — rather than the frame.

That's what makes the record both affordable and useful at the same time. Every question in the list at the top of this post is a question about episodes, not about frames. And with the raw footage retained behind it as evidence, retrievable by reference, "watched" stops meaning a person was looking and starts meaning the record can be asked.

The asset is stranded only until asking gets cheap#

Markets have started pricing the gap. Surveillance infrastructure — cameras, storage, VSaaS plumbing — is forecast to grow at roughly 9–12% a year. The analytics layer on top of it is forecast at roughly 28–33%, with individual firms putting AI video analytics anywhere from $32B to $133B by the end of the decade. Scopes differ enough between research firms that the right way to read those numbers is as a range and a direction, not a point estimate. The direction is consistent: value is migrating from the sensor to the understanding.

Which is the market's way of saying what this post says in words. The cameras are a sunk, appreciating archive of unanswered questions. They stay stranded only until the cost of asking falls below the value of the answer.

Where this argument is weakest#

We'd rather state this than have you find it.

The headline trio is vendor framing. 1.5B cameras / 7T hours / <1% watched originate in category marketing. We use them because they're the industry consensus and directionally robust, but no independent census of enterprise cameras exists. If the real installed base were half that, the argument weakens in degree, not in kind.

The 90% figure is ours. It matches the public price structure of every component, but it has not been independently published. The argument survives it being 70% or 95% — what matters is that the dominant cost center is also the one on the steepest downward curve.

Storage bands are planning estimates. Per-camera vector rates vary with scene, mounting, and write policy. The ~20× naive-versus-gated ratio is modeled, not measured across a production fleet, and anyone deploying this should validate it on representative feeds rather than trusting our band.

videoembeddingsmarketevent-drivenvector search
Less Than 1% of the World's Video Is Ever Watched | Miriel