The Focoos VLM
A model that reads
what is happening.
Ours. Small enough to run in your cabinet, trained on the acts that matter, and it answers in words.
Where it comes from
We make models
that fit on one machine.
Focoos AI came out of a computer vision lab, and efficiency is what the team does: the same accuracy, on a fraction of the hardware. That is why the model fits in your cabinet instead of taking your video away.
Who we areFounded out of Politecnico di Torino, with more than 400 published papers behind the people who build it.
Every version is made to run on smaller hardware, not bigger. That is the metric we chase.
The model is ours, trained by us. It is not somebody else's API with our name on it.
What it is
A model that sees,
and then writes.
A VLM is a vision language model: pictures go in, words come out. Ours has been trained on surveillance footage and asked one question, over and over, on every camera: is anything happening here. It works the way a good operator does.
Streams in, a sentence out. No rules, no zones, no thresholds to tune.
One model, seventeen behaviours
What it recognises,
and why it is one model.
Trained on footage of the acts that matter, not on ordinary street video.
One model, not one per problem
A forecourt, a loading bay, a kerb. Nothing gets rebuilt for a new scene.
Nothing to train
It arrives ready. No labelling, no learning period, nothing to tune per camera.
It fits in one machine
Small enough for one cabinet. That is what keeps your video on your site.
It shows its working
Every answer carries the sentence behind it, so an operator can disagree.
What it recognises today
Seventeen behaviours, the same list for every customer, and the list grows with every version of the model. Click one.
Always improving
It never stops
getting better.
You never train it: we do. The model in a year will be better than today's, and nothing on your site has to be rebuilt for it. Every confirm and every dismiss your operators record is a signal, and the whole fleet gets each improvement at once.
An operator settles it
Confirmed, or dismissed. Either way it is recorded as a signal.
The hard cases come back
The ones it was unsure about are the ones worth learning from.
The model is retrained
On real footage from real sites, never on a public dataset.
Every Box gets it
The update ships out to every Box in the fleet, all at once.