The Focoos VLM

A model that reads
what is happening.

Ours. Small enough to run in your cabinet, trained on the acts that matter, and it answers in words.

Where it comes from

We make models
that fit on one machine.

Focoos AI came out of a computer vision lab, and efficiency is what the team does: the same accuracy, on a fraction of the hardware. That is why the model fits in your cabinet instead of taking your video away.

Who we are
A research team

Founded out of Politecnico di Torino, with more than 400 published papers behind the people who build it.

Efficiency first

Every version is made to run on smaller hardware, not bigger. That is the metric we chase.

Built here

The model is ours, trained by us. It is not somebody else's API with our name on it.

What it is

A model that sees,
and then writes.

A VLM is a vision language model: pictures go in, words come out. Ours has been trained on surveillance footage and asked one question, over and over, on every camera: is anything happening here. It works the way a good operator does.

the stream, all of itstreamsFOCOOS VLMone model, on your BoxDescriptionwhat it saw, in a sentenceAlarmyes, or noTypewhich conditions fired

Streams in, a sentence out. No rules, no zones, no thresholds to tune.

One model, seventeen behaviours

What it recognises,
and why it is one model.

Trained on footage of the acts that matter, not on ordinary street video.

01

One model, not one per problem

A forecourt, a loading bay, a kerb. Nothing gets rebuilt for a new scene.

02

Nothing to train

It arrives ready. No labelling, no learning period, nothing to tune per camera.

03

It fits in one machine

Small enough for one cabinet. That is what keeps your video on your site.

04

It shows its working

Every answer carries the sentence behind it, so an operator can disagree.

What it recognises today

Seventeen behaviours, the same list for every customer, and the list grows with every version of the model. Click one.

Always improving

It never stops
getting better.

You never train it: we do. The model in a year will be better than today's, and nothing on your site has to be rebuilt for it. Every confirm and every dismiss your operators record is a signal, and the whole fleet gets each improvement at once.

1

An operator settles it

Confirmed, or dismissed. Either way it is recorded as a signal.

2

The hard cases come back

The ones it was unsure about are the ones worth learning from.

3

The model is retrained

On real footage from real sites, never on a public dataset.

4

Every Box gets it

The update ships out to every Box in the fleet, all at once.