The brain, on your own network.

It watches every camera you already own, understands what is happening and decides there. It records too.

The core

Focoos VLM

A vision language model of our own, small enough to run in your cabinet. It is the reason the device understands what is happening instead of detecting that something moved.

Three things make it strong, and they are the three things nobody can copy quickly: who built it, what it is optimised for, and what it is trained on.

Built by a research team

Focoos AI came out of a computer vision lab, with more than four hundred published papers behind the people who build the model. That is who is training it, not a vendor reselling an API.

Made smaller, on purpose

Every version is made to run on smaller hardware, not bigger. That is the metric we chase, and it is the only reason the model can sit in your cabinet instead of in somebody’s cloud.

Trained on what actually happens

Trained on surveillance footage, on the acts that matter, which is why it reads a forecourt at three in the morning and a general model does not.

A model that sees, and then writes.

A VLM is a vision language model: video goes in, words come out. Ours is trained on surveillance footage, and on each camera it carries one question per event type you have switched on, each asked of the same scene. It works the way a good operator does: it watches, and then it says what it saw.

the stream, all of itstreamsFOCOOS VLMone model, on your deviceDescriptionwhat it saw, in a sentenceAlarmyes, or noTypewhich conditions fired

Streams in, a sentence out. No rules, no zones, no thresholds to tune.

Twenty event types. Each camera gets its own set.

This is not a detector putting labels on objects in a frame. Each of the twenty is something that happens: an act, read over time, the way a person reads it. One model covers all of them, and on each scene you switch on the ones that belong there.

Damage and vandalism

7

who breaks, forces, defaces or tampers with something

Suspicious behaviour

6

the signals that come before an attack, or an act of vandalism

Perimeter and entry points

2

who goes where they should not, or throws something in

Violence and emergencies

3

when one or more people are in danger

Presence

2

presences nobody wants, and absences nobody explained

Four steps, and no rules to write.

This is everything it takes to put a camera under analysis. No zones to draw, no thresholds to tune, no model to train.

01

Add the camera

The RTSP stream of the camera you already have, H.264 or H.265. Nothing to replace.

02

Choose whether to record

How many days the recordings are kept on the device, camera by camera.

03

Choose what to switch on

The point of view sets the usual event types for that spot. For example, on an ATM: tampering, unusual objects, damage, covered face.

04

Switch it on

From that moment the camera is under analysis. There is no fourth week of tuning.

What it looks like on site.

The device has a console of its own, reachable on your network and working with the internet down: the videowall of the scenes it watches, the setup behind them, and the events it has raised. These are captures of the running product.

The numbered points explain what you are looking at.

How much machine you need.

Everything up to here is one piece of software, and it is the same piece on every machine: the same model, the same twenty event types, the same four steps. What the machine changes is one thing only, how many cameras it carries. Past that you put machines together in whatever combination suits the site.

124487296
MachineClassCameras per machine

Figures are estimates and depend on resolution and the scene.

Every clip on this site is real surveillance footage, not a render, and every console screen is a capture of the running product. The footage comes from third-party sources: copyright remains with its respective owners and it is shown here only to illustrate what the model reads.

© 2026 Focoos AI · Torino, Italy focoos.ai Contact Privacy