PerspectiveSeptember 2, 2026·7 min read

Four Cost Curves Are Collapsing at Once, and They All Point at Video

Four Cost Curves Are Collapsing at Once, and They All Point at Video

Watching the world was always desirable. It is becoming affordable — and not on one axis, which would be a trend, but on four at once, which is a phase change.

This post prices each of the four, presents the counter-evidence we think is strongest, and then does the arithmetic that matters: when does continuously understanding one camera cross under a budget line you'd actually approve?

First, where the money goes#

A real-time video pipeline has six cost centers. Decode the stream. Detect and track what's in frame. Embed what matters into vectors. Store them. Search them. Answer questions over the results.

They are not remotely equal. In our operational experience, embedding generation has been ~90% of total pipeline cost. That figure is first-party — we haven't published a methodology for it the way we have for our throughput benchmarks — and we label it as such everywhere it appears.

The structural reason is duty cycle. Embedding inference runs continuously, on every stream, whether or not anyone ever asks a question. Storage and search amortize across queries and are fast in absolute terms: in our production retrieval stack, graph traversal runs 5–19 ms and vector search 50–120 ms, against LLM answer synthesis at 7–10 seconds. Retrieval is never the bottleneck. The meter that runs all day is the embedder.

Which means: whoever makes embeddings cheap makes the whole category viable. Every curve below matters, and the embedding curves matter most.

Collapse one: inference, roughly 10× per year#

The best-documented price collapse in the history of computing is happening in model inference, and it is happening at constant capability — which is the only comparison that means anything.

GPT-3-equivalent capability cost $60 per million tokens in November 2021 and $0.06 per million by November 2024. That's 1,000× in three years. Epoch AI, measuring price-to-hit-a-fixed-benchmark across many benchmarks, finds declines between 9× and 900× per year depending on the benchmark, with GPT-4-level performance on PhD-science questions falling about 40×/year. GPT-4-class output that sold for $20–60 per million tokens in 2022–23 sells for around $0.40 today.

Dedicated embedding models ride the same slope. OpenAI's ada-002 at $0.10 per million tokens in 2022 became text-embedding-3-small at $0.02 in 2024, and $0.01 via batch — a 10× effective decline in about two years — while open-source embedding models push the marginal cost down to raw GPU time.

Nobody should extrapolate 10×/year forever, and we don't. The 2026 consensus view is 3–5× per year through 2027 before tapering, and that's the rate we plan against below.

Collapse two: hardware, with a step change underneath it#

Under the model curve sits a hardware curve. GPU FLOP/s per dollar has doubled roughly every 2.2–2.5 years, with accelerator performance-per-dollar improving 30–40% a year across 2012–2025.

On top of that trend came a step. Inference providers report Blackwell delivering up to 10× lower cost per token than Hopper in production, and SemiAnalysis estimates up to 35× for agentic workloads on GB300 NVL72.

For video specifically, the per-stream arithmetic already lands at pennies per camera-hour. An L40S- or 4090-class card handles 8–16 concurrent analytics streams with NVDEC decode and TensorRT batching. Our own published ingest numbers sit in the same envelope: ~8,000 vectors/second on a single four-L4 node, scaling near-linearly across nodes. At the edge, dollars per TOPS fell about 10× in seven years — Jetson Xavier at $1,299 for ~30 TOPS in 2018 against Orin Nano Super at $249 for 67 TOPS in December 2025.

That last one has a twist, and it's in the counter-evidence section below.

Collapse three: video understanding is already cents per hour#

Here is the number that would have sounded absurd in 2022. Frontier-lab pricing now puts an hour of video understood at $0.04–0.33, depending on model tier and resolution.

Tier$ / hour of videoNote
Flash-Lite$0.04–0.10low-resolution tokenization
Flash, low-res + batch≈$0.11
Flash, default≈$0.33~300 tokens/s of video
Pro≈$1.37a 10×+ spread from routing alone

Published Gemini list pricing, mid-2026. Batch and low-resolution discounts as published.

Operational levers compound on top of the list price: caching (70–90% savings on repeated content), batching (~50%), self-hosting (5–20×). One documented video-AI workload fell from ~$700K/month to ~$200–250K/month before any hardware refresh.

Collapse four: vector storage, 4–32× and a broken RAM wall#

The historical objection to "vectorize everything" was that vector indexes lived in RAM, and RAM costs 10–50× what SSD costs per GB.

That wall is broken. Disk-based indexes in the DiskANN/Vamana family serve 1B+ vectors on a single 64 GB-RAM machine at 95%+ recall with sub-5 ms search; Kioxia's AiSAQ demonstrated 4.8B vectors on one server. Quantization multiplies the effect — int8 gives 4×, binary up to 32× — and one documented example took 100M vectors from $1,459/month to about $45/month. GPU acceleration went mainstream in 2025–26: AWS OpenSearch reports 9.3× faster billion-scale index builds at 3.75× lower cost with cuVS, Elasticsearch up to 12× indexing throughput, GPU DiskANN builds 40×+ over CPU.

Trillion-scale is now a marketing claim in this category, which is worth saying plainly. The differentiated evidence is what somebody has actually run. Ours is the one we published in June: 100 billion encrypted vectors searched at a flat 199 ms p50 on 20 commodity CPU boxes, with median latency moving about 21 ms across a 20× increase in dataset size. Encrypted the whole way — that flat number is the encrypted number.

The counter-evidence we take seriously#

A cost-curve argument that only cites declines is marketing. Three things cut the other way.

Curves are not monotonic. In July 2026 NVIDIA raised Jetson prices by up to 101% — Orin Nano Super from $249 to $399, AGX Thor to $5,499 — amid AI-boom component demand. The seven-year edge trend is down ~10×; the last twelve months moved hard against it. In the near term this actually strengthens the case for centralizing the expensive embedding lane, since shared datacenter GPUs amortize better than per-site edge boxes. But it's a real reminder that a supply shock can pause any of these curves for quarters at a time.

Demand is booming against constrained supply. The same boom driving price per token down is bidding the underlying silicon up. Net direction at the inference-price layer has stayed down, but 3–5×/yr is the honest planning rate, not 10×.

"Cheaper" is not "free." At fleet scale the record still costs real money. The claim isn't that cost goes to zero. It's that cost crosses below the value of the answers — and below the labor it displaces — and that the crossing has a date.

The arithmetic, with every assumption visible#

Take the published $0.33/hour tier. Assume a 3×/year decline (the conservative end of consensus). Assume a 30% duty cycle, meaning event-gated embedding skips 70% of the hours outright.

code
cost(year) = baseline $/hr × 730 hrs/mo × duty ÷ rate^(year − 2026)

That camera costs about $72/month to understand in 2026, crosses under $10/month in 2028, and reaches roughly $0.90/month by 2030. For comparison, human remote video monitoring — the closest thing to a market price for continuous watching — runs $50–150 per camera per month today.

Set the dials pessimistically and the conclusion bends without breaking. At 2×/year the crossing moves to 2029. At the cheap-tier baseline of $0.04/hour it is under $10 from the start with no decline at all. Only the combination of a full-price baseline and a sub-2× decline pushes the crossing past 2030.

The variable that moves the date most is duty cycle — and duty cycle is an engineering choice, not a market condition. Event-gated pipelines are how you pull the date in instead of waiting for it.

Two more things about that model, both of which make it conservative. It prices frontier-API video understanding; self-hosted embedding at fleet scale sits below the API curve, which is why our own infrastructure math lands lower. And it prices only the questions that need a model at all. Most questions about a video record have a schema and compile to SQL — they cost database time, not model time, which is a separate post and a bigger deal than any curve here.

What we take from this#

The window where continuous video understanding becomes ordinary is roughly 2027–2030, and the argument doesn't depend on any single curve holding. Inference could stall at 2×/year and vector storage alone would carry it. Jetson pricing could stay elevated and centralized architectures would absorb it. What doesn't have a curve working in its favor is the human on the monitoring seat.

We build for the far side of that crossing: encrypted vectors, event-gated writes, and a query surface where the expensive model is the last resort rather than the first. Miriel is a member of NVIDIA Inception, and a good deal of the above is why — the hardware curve and the storage curve are the two we most need to keep riding.

videoembeddingsGPUcostvector searchNVidia
Four Cost Curves Are Collapsing at Once | Miriel