PRODUCT — PERCEPTION

Give your AI a memory of the real world

Point a camera, a microphone, a sensor or a folder of PDFs at it, and what arrives becomes rows you can query, search semantically and alert on. Months of it, on one timeline, answerable in a sentence.

// if it streams, lands, or arrives — it ingests it

perceptdb · consolevideo · listening
VideoAudioSensors
natural language → sql + vector · evidence returned with every answer
WHAT IT IS

A perception database: everything becomes a row, and heavy bytes are always represented by one.

Video, audio, logs, chat, sensor feeds, market ticks, phone-call transcripts, PDFs. Every source flows through the same pipeline and lands in the same relational surface. A file becomes an object — bytes in object storage, a row in Postgres. A log line, a chat message or a sensor reading becomes a stream event. Embeddings are not a side-car index; they are rows too, attached to the exact data they describe.

Perception workers then pull from a queue and run models over anything new — scene understanding, CLIP, detection and tracking, transcription, text embedding — and write their findings back as more rows. Captions, detections, tracks, object events, embeddings. All of it queryable through one SQL surface, where a semantic-similarity function splices vector search into ordinary SQL, and alert rules run continuously over the same rows.

The consequence that matters: the store is the memory. Questions are answered from the database, not from a process that had to have been watching at the time. That is what makes an answer survive a restart, and what makes an audit trail possible a year later.

WHAT A FRAME BECOMES
What a frame becomesA camera frame and three other sources fan out into detections, text, embeddings, transcripts and sensor readings. All of them land on one shared timeline, which a single question sweeps across to return an answer with the evidence frames attached.ONE SAMPLED FRAMEperson 0.94torch 0.88cam-3 · 14:02:11 · 1.82 mAND EVERYTHING ELSERTSP · depthmicrophonesPLC historianPDF corpusOpen-vocabulary detection: theclasses are text prompts, so a newobject needs no retraining and nolabelled data.detections · trackstext / OCRembeddingstranscript · voicessensor readingsWritten at boundaries — whensomething appears, changes, movesor leaves — not once per frame.≈20× less storageONE TIMELINEThe store is the memory. Answerssurvive a restart; an audit trailis queryable a year later.“who had it last?”evidenceframes
Per sampled frame the pipeline runs detection, embedding, OCR, identity and window-level description in parallel. Those outputs, and everything arriving from every other source, land on one shared timeline.

Open-vocabulary detection

The classes are text prompts, not a fixed label set. Type "welding hood" or "chain hoist" and the cameras start finding it — prompts are re-embedded at runtime, so there is no retraining and no labelled data.

Whole-frame and per-object embeddings

In CLIP space, so a frame is retrievable by a description of what was happening in it rather than by a class name someone chose in advance.

Text and OCR

A placard, a heat-number stamp, a job traveler on a cart — searchable as text.

Faces, voices and a concept grid

Face embeddings and speaker embeddings, clustered rather than matched against a roster.

Sliding-window clip embeddings

So motion is retrievable, not just appearance — plus a small vision-language model that writes one sentence per window, which is what lets an answer say what it found instead of only pointing at it.

THREE DECISIONS THAT MATTER MORE THAN THE MODELS
  • Write at boundaries, not per frame. Vectors are written when something appears, changes, moves or leaves — roughly 20x less storage than naive detection logging. That is the difference between a multi-camera install staying affordable over months and over days.
  • Metres, not pixels. With depth cameras calibrated into one world frame, detections become measured world coordinates with real object width, height and thickness read from the depth region inside each box. Depth frames fuse into a coloured surface mesh of the actual space, so “where” has an answer rather than a bounding box. Several cameras feed one floor plan, and counts merge by per-label maximum — never sum — because two lenses on the same bay see the same person, and “at least N” is honest where 2N would be a lie.
  • Identity that starts anonymous. Face sightings cluster by embedding similarity and begin as “face 3”. Naming one is retroactive: every episode since first sighting transfers to that name. Voices cluster the same way and link to a person by noticing that voice is heard while exactly one named person is alone on camera. Thresholds are deliberately conservative — the system would rather show you an honest “face 4” than confidently call a stranger by the wrong name.
MEASURED, NOT ESTIMATED
7,981vectors/second through the full ingest path, one node, 4x NVIDIA L4
18,175vectors/second across three nodes — linear, no cross-node coordination
~20xless storage than logging every detection on every frame
< 6 GBVRAM for the entire GPU inference tier

Single-node figure measured on one g6.12xlarge with 16 workers pinned four-per-GPU in fp16, running the complete path — frame sampling, whole-frame and per-object embedding, open-vocabulary detection, OCR, concept indexing and face embedding — on busy footage; quiet footage runs about 2,900 vectors/second. The inference tier is four services that together fit under 6 GB of VRAM, so any 16 GB card is comfortable and pipeline workers stay thin enough to pack densely.

BUILT ON NVIDIA
NVIDIA Inception Program

Perception is an NVIDIA-shaped problem, and we build it on NVIDIA at every tier.

  • At the sensor. Orbbec Femto Mega depth cameras run their depth engines on an onboard NVIDIA Jetson, streaming processed depth alongside 4K RGB over PoE — so the geometry of a space is measured at the sensor rather than reconstructed after the fact.
  • In the index. Recorded clips are captioned and embedded with NVIDIA Cosmos, turning what a space recorded into searchable memory — retrievable by what happened, not just by what was in frame.
  • Under the vector store. The vector database builds its indexes on NVIDIA cuVS — coarse IVF centroid training and vector-to-list assignment run on cuVS, and builds stay GPU-resident. Vectors and their metadata are encrypted before they enter that index, and nearest-neighbour search runs against the encrypted index rather than a decrypted copy of it.
  • In the architecture. The result is built to the same shape as the NVIDIA Metropolis VSS blueprint: we consume frame-level understanding and clip summarization, and add the layer above it — months of cross-modal memory, the join to the governing document, and an audit trail queryable later.
  • And in the company. Miriel is a member of the NVIDIA Inception program.

In September 2026 we wired a house in Palo Alto into a single perception system and ran two days of live, interactive demos out of it — Jetson-based depth cameras, one encrypted index, and a house that could answer questions about its own past. No slides and no recordings, which is still the only demo we think is worth giving.

WHERE IT RUNS

Cloud, your VPC, or on-premises

The same stack in all three. The GPU inference tier is a single host running four services that together need under 6 GB of VRAM.

At the edge

Camera-side inference on the floor, with only events and embeddings crossing the network — the natural shape when bandwidth is constrained or the footage should not leave the building.

Scoped by project

Resources carry per-user grants, and the knowledge graph inherits those boundaries rather than creating new exposure. Cross-document insight only emerges from material the asker could already see.

WE CALL IT PERCEPT

The perception layer has a name — Percept — and it is the database underneath everything on this page: vectors, columnar metadata, time series and a relationship graph in one place, so a single question resolves across documents, video, audio and sensor readings at the same time with the evidence attached.

Most teams never need to think about it that way. You connect a source and ask a question. But when someone on your side asks what is actually holding all of this, that is the answer.

GET STARTED

We would rather show you than tell you.

There is a working multi-camera installation doing all of this, and running it live is more convincing than a diagram.

// no spam, just a beta waitlist