Every video-AI demo has the same shape: point a model at a stream, watch labels appear over the picture, be impressed. Every video-AI fleet runs into the same wall about six weeks later, and it is not a model wall. It is that the demo's write policy, multiplied by 500 cameras and a 90-day retention window, produces a storage bill nobody approved.
This post is about the four decisions that separate the two, starting with the one that matters most.
The naive policy is wrong by about 20×#
Analyze at 5–15 fps with several detections per frame, write a vector for each, and a single camera persists billions of records a year. Nearly all of them are repeats: the same parked forklift, the same shelf, the same empty corridor, frame after frame after frame. You are paying full price to store the fact that nothing changed.
Modeled against gated policy on the same footage, naive every-detection-is-a-write overstates real volume by roughly 20×. That's a planning band from our own scale modeling, not a production measurement, and it's the single number we most want people to validate on their own feeds before they size anything.
The fix is to write on events, not frames. Detection and tracking still run continuously — you can't gate what you haven't looked at — but a vector is persisted only at meaningful boundaries:
- a track enters the scene
- an object's appearance changes materially
- a better-quality view replaces a worse one (the object turned toward the camera, or moved out of glare)
- a periodic refresh, so long dwells stay searchable
Under that policy, a camera writes on the order of 0.1–10 vectors per second depending on scene density — a quiet perimeter fence at the bottom of that range, a crowded dock face at the top. And the queryable unit stops being the frame and becomes the episode: something appeared, persisted, moved, left.
That last part is not a storage optimization; it's the reason the record is useful. "When did the pallet in bay 12 leave?" is a question about an episode. Nobody has ever wanted a frame.
The pipeline is a dial, not a fixed cost#
The second decision is refusing to ship one pipeline. Ingestion is assembled per deployment, per zone, and where it matters per camera — and every stage is a quality/cost/speed dial.
flowchart LR A[Camera / RTSP] --> B[Event gate<br/>motion, schedule] B --> C[Detect + track<br/>YOLO-class, crops] C --> D[Tile + embed<br/>whole frame, squares] C --> E[OCR<br/>plates, labels, badges] C --> F[Clip encoder<br/>768-d or 256-d] C --> G[Faces<br/>opt-in, FHE path only] D --> H[Data layer] E --> H F --> H G --> H H --> I[Vectors, encrypted] H --> J[SQL tables] H --> K[Time-series + graph] H --> L[Object storage<br/>clips, crops, originals]
A stream can be event-gated or continuous. Frames can be embedded whole or tiled into squares when small-object recall matters. A detector can run with or without re-embedded crops. OCR can be on for plates, badges, and labels, or off entirely. Face embedding is opt-in per zone and rides the homomorphic-encryption path when it's on. The clip lane can run a 448p/768-d quality tier or a 224p/256-d compact tier.
Dimensionality alone moves storage linearly: at FP16, bytes per vector is just
dimension × 2, so a compact tier is a 4.5× storage difference against a
1,152-d native tier before anything else changes. That is how the same
architecture serves a quiet perimeter fence at cents per month and a dense dock
face at full fidelity. You tune the pipeline to the question the zone needs to
answer, rather than tuning the question to the pipeline.
The same dial is the accuracy lever. Speed is only one of two gates on a real deployment; the other is model quality, and a stock detector that's right for a dock face is wrong for body-cam footage or a residential doorway. So the model slot gets treated like every other stage: any model — including anything on Hugging Face — can be registered, deployed, and benchmarked in place, with per-model throughput measured in production and a dynamic allocator shifting GPU share between embedding lanes as load moves. Compact models where speed and cost dominate; the highest-accuracy model exactly where the question demands it; swapped per customer without touching the rest of the pipeline.
Edge and cloud are one record in two tiers#
The third decision is about where the dial sits physically.
A ~$300 Jetson-class box on premises runs compact detectors and embeddings for instant local decisions, and keeps raw video off the wire. A datacenter-class GPU runs the flagship models and deep search. Different models and different dimensionalities per tier, chosen by what each tier is for.
The part that makes this work rather than merely sound good: both tiers run CyborgDB, and the data layer replicates between them. Two instances, one record. A query doesn't care which tier a vector was born on.
This also happens to be the right hedge on hardware pricing. Edge silicon got about 10× cheaper per TOPS over seven years and then, in July 2026, got up to 101% more expensive in a single price revision. An architecture that can move the expensive embedding lane between tiers absorbs that; one that assumed a $249 edge box does not.
Pin the model contract, or the record expires#
The fourth decision is the boring one that saves you in year three.
Every vector space is pinned to model, version, normalization, and dimension. Compressed tiers are adopted only after passing a recall gate against the model-native dimension — you don't get to ship a 256-d tier because it's cheaper, you ship it because it held recall on your footage.
Fleets outlive models. A camera installed this year will still be recording when the embedding model you'd choose today is three generations stale, and the question "is a 2026 vector comparable to a 2029 vector" needs to have an answer that isn't a shrug. The contract is what makes a five-year record queryable in year five, and re-embedding on a model change is a migration you plan rather than a surprise you discover.
The same discipline extends to what a lane is allowed to be used for. Re-identification features, for instance, are scoped to track association only, never to general semantic search. That's a privacy boundary as much as an engineering one, and encoding it in the contract is how it survives the next feature request.
What it adds up to#
Put the four together and fleet economics stop being mysterious.
Take 500 cameras averaging about two activity events per minute, each event analyzed into ~105 vectors (roughly 5 analyzed frames × ~20 vectors each: whole frame, objects, text), plus a 5% non-camera sensor stream — badges, RFID, GPS, audio, alarms all embed too, as compact ~256-d vectors. Retain half of it for 90 days, let the rest expire after a 2-day buffer. 768-d object vectors, one clip embedding per 10 seconds.
That fleet runs roughly $18 per camera per month on infrastructure. For comparison, human remote video monitoring runs $50–150 per camera per month to watch a fraction as well.
Every number in that paragraph is a dial, and the interesting thing is which dials move the bill. Dimensionality moves it linearly. Retention share moves it linearly. Scene density moves it linearly. Event gating moves it by the ~20× factor at the top of this post — which is to say the write policy is worth more than every other lever combined.
One honest wrinkle: the bill moves in whole-node steps. A small fleet sits on a one-ingest-plus-one-storage-node floor no matter how the dials move, so a 65-camera deployment pays a real per-camera premium over a 5,000-camera one for exactly the same architecture. When someone quotes you a per-camera rate, ask which of the two they're quoting.
Where we'd want to be checked#
The per-camera vector rates are planning bands, not fleet measurements. Scene, mounting height, and write policy all move them, and we expose them as inputs rather than constants because we don't think a single number is honest.
The ~8,000 vectors/second/node ingest figure is measured, on 4× L4 nodes with busy footage, and published with methodology. The other node profiles we quote — L40S, single-L4, Jetson edge — are scaled estimates from that measurement, and we label them that way.
And the 20× gated-versus-naive ratio is modeled. It's the load-bearing number in this entire post, we haven't measured it across a production fleet, and it is the first thing we'd re-derive on your footage rather than ask you to take on faith.


