Damage and vandalism
7who breaks, forces, defaces or tampers with something
It watches every camera you already own, understands what is happening and decides there. It records too.
The core
A vision language model of our own, small enough to run in your cabinet. It is the reason the device understands what is happening instead of detecting that something moved.
Three things make it strong, and they are the three things nobody can copy quickly: who built it, what it is optimised for, and what it is trained on.
Focoos AI came out of a computer vision lab, with more than four hundred published papers behind the people who build the model. That is who is training it, not a vendor reselling an API.
Every version is made to run on smaller hardware, not bigger. That is the metric we chase, and it is the only reason the model can sit in your cabinet instead of in somebody’s cloud.
Trained on surveillance footage, on the acts that matter, which is why it reads a forecourt at three in the morning and a general model does not.
A VLM is a vision language model: video goes in, words come out. Ours is trained on surveillance footage, and on each camera it carries one question per event type you have switched on, each asked of the same scene. It works the way a good operator does: it watches, and then it says what it saw.
Streams in, a sentence out. No rules, no zones, no thresholds to tune.
This is not a detector putting labels on objects in a frame. Each of the twenty is something that happens: an act, read over time, the way a person reads it. One model covers all of them, and on each scene you switch on the ones that belong there.
who breaks, forces, defaces or tampers with something
the signals that come before an attack, or an act of vandalism
who goes where they should not, or throws something in
when one or more people are in danger
presences nobody wants, and absences nobody explained
This is everything it takes to put a camera under analysis. No zones to draw, no thresholds to tune, no model to train.
The RTSP stream of the camera you already have, H.264 or H.265. Nothing to replace.
How many days the recordings are kept on the device, camera by camera.
The point of view sets the usual event types for that spot. For example, on an ATM: tampering, unusual objects, damage, covered face.
From that moment the camera is under analysis. There is no fourth week of tuning.
The device has a console of its own, reachable on your network and working with the internet down: the videowall of the scenes it watches, the setup behind them, and the events it has raised. These are captures of the running product.
The numbered points explain what you are looking at.
Everything up to here is one piece of software, and it is the same piece on every machine: the same model, the same twenty event types, the same four steps. What the machine changes is one thing only, how many cameras it carries. Past that you put machines together in whatever combination suits the site.
| Machine | Class | Cameras per machine |
|---|
Figures are estimates and depend on resolution and the scene.
This is the shape of the product, not a setting you switch on afterwards. The device opens one connection towards the cloud and nothing can reach back down it: there is no inbound port to open, and no rule for your IT to argue about. What is not sent cannot leak.
The alarm: what the model saw, its type, and a clip of the moment it happened.
The continuous recording, and the whole archive on the disk.
Live and playback, over the same connection, only while an operator is watching.
GDPR today, the AI Act and the Cyber Resilience Act as they land.
The footage never leaves the site that filmed it. What crosses is one alarm and its clip, kept for as long as you set.
Acts, not identities. No face recognition, no biometric identification, which is where the heavy obligations sit.
No inbound port, one connection going out. The formal work is under way ahead of the dates.
"*" indicates required fields
We’ll get back to you shortly