Milestone Systems has unveiled an innovative vision language model (VLM) designed to enhance traffic understanding, making use of NVIDIA Cosmos Reason's capabilities.
This release includes two major products: a Video Summarisation tool specifically for XProtect® Video Management Software and a VLM as a Service designed for third-party integrations. The Video Summarisation tool streamlines the process of searching through visual data and automating reports.
Video summarisation tool
Video Summarisation tool efficiently processes camera footage and provides descriptive text summaries
Video systems generate vast quantities of data, and the manual review of this footage is both time-intensive and labour-intensive.
Now, with the introduction of Milestone Systems’ Video Summarisation tool - a generative AI-powered addition to the XProtect Smart Client - users can automate workflows, dramatically saving time and reducing false alarm fatigue by as much as 30%, according to initial findings.
The Video Summarisation tool efficiently processes camera footage and provides descriptive text summaries. Operators input a video snippet along with a specific prompt, and the model produces a textual summary rapidly.
Key capabilities
- Transform video clips into structured text summaries directly within the XProtect Smart Client.
- Enable searching of summaries based on visual content rather than relying on timestamps or tagging.
- Bookmark and filter summaries to optimise review processes.
- Integrate seamlessly with existing XProtect event and rule systems to trigger automatic summaries for select alerts or alarms.
- Concentrate on valid events by excluding unnecessary motion or noise.
- Access region-specific VLMs starting from the US and EU, with plans to expand further.
This tool is freely downloadable, with a straightforward installation in the XProtect Smart Client, and costs are incurred only when utilising the VLM.
VLM as a service for developers
Milestone’s Hafnia VLM as a Service (VLMaaS) grants developers, integrators, and partners access to advanced video intelligence through NVIDIA technology, facilitating the creation of AI-driven solutions without the need for extensive AI system management or fine-tuning.
VLMaaS allows developers to expedite AI and analytics development, reportedly reducing the effort
This service streamlines the incorporation of sophisticated video intelligence features for application development, whether for launching a new product or expanding an existing platform.
VLMaaS allows developers to expedite AI and analytics development, reportedly reducing the effort necessary for fine-tuning a VLM model by up to 70 times compared to traditional methods.
Key capabilities
- Utilise a highly accurate vision language model optimised for traffic data, built on NVIDIA Cosmos Reason.
- Support traffic-related operations with prompt-based instructions.
- Simplify integration through API-first delivery over HTTPS.
- Employ models fine-tuned for the US and EU, with additional regions to be added.
- Designed for standalone solution development or integration with Milestone’s product offerings.
- Utilise responsibly sourced training data with auditable lineage, compliant with GDPR and EU AI Act standards for model fine-tuning.
Pricing for VLMaaS is based on API usage, eliminating the need for significant upfront investments or customised training costs.
Vision Language Model as a Service
Andrew Burnett, Acting Chief Technology Officer at Milestone Systems, stated, "With the Vision Language Model as a Service and Video Summarisation for XProtect, we're tackling some of the most challenging bottlenecks: video overload and time-consuming manual work. Operators get immediate insight directly within XProtect; builders get API‑first access to production‑ready intelligence without bespoke training or heavy infrastructure.”
Customers in cities such as Genoa, Italy, and Dubuque, Iowa, US, are eager to implement these capabilities to enhance traffic management systems with advanced video intelligence solutions.
Responsible AI and real-world data
The new solutions are supported by Milestone’s Hafnia VLM, refined using 75,000 hours of real-world video data from Europe or the US, managed with NVIDIA Cosmos Curator, and operational either in cloud settings or regional data centres.
Integrating NVIDIA Cosmos Reason VLM with Milestone’s data positions this as a top-tier video AI platform.
Milestone Systems, a world pioneer in data-driven video technology, released an advanced vision language model (VLM) specialising in traffic understanding and powered by NVIDIA Cosmos Reason.
The VLM powers two new products: a Video Summarisation tool for XProtect® Video Management Software and a VLM as a Service for third-party integrations. Video Summarisation for XProtect allows users to search summaries from visual data and automates reporting.
Video Summarisation tool
Video systems capture vast amounts of data, and reviewing footage remains time consuming and largely manual. With Milestone Systems’ new Video Summarisation tool – a generative AI-powered plug-in for the XProtect Smart Client – users and operators can now rely on a specialised product that automates operator workflows, saves valuable time, and reduces false alarm fatigue significantly. Early reports show video summarisation could reduce operator false alarm fatigue by up to 30 %.
The Video Summarisation tool analyses camera footage and describes what's happening. Users simply send a snippet of video and a prompt describing their request, and the model will generate a text summary in seconds.
Key capabilities
- Convert video segments into structured text summaries inside XProtect Smart Client
- Search summaries based on video content, rather than timestamps or manual tagging
- Bookmark and filter summaries to streamline review workflows
- Integrate seamlessly with existing XProtect event and rule logic to trigger automated summaries based on specific alarms or alerts
- Focus attention on valid events by filtering out irrelevant motion or noise
- Access customised, sovereign VLM’s per region, starting with the US and EU. More regions to follow.
The Video Summarisation is free to download and takes only a few minutes to install directly in the XProtect Smart Client. And users only pay when prompting the VLM.
VLM as a Service for developers: Add production-ready video intelligence
With Milestone’s Hafnia VLM as a Service (VLMaaS), developers, integrators and partners get API access to production-ready video intelligence built on NVIDIA’s latest technology and fine-tuned on responsibly sourced data.
The VLMaaS helps developers create AI-powered solutions quickly without needing to set up, fine-tune or manage their own AI systems – it enhances any existing solutions with generative AI, regardless of the level of analytics currently in place. This makes it fast and simple to add advanced video intelligence features to applications, whether it’s testing a minimum viable product (MVP) or scaling a platform.
With VLMaaS, the development of AI and analytics can be accelerated significantly – up to 70 times less effort than doing the work to fine-tune a VLM model to do the same.
Key capabilities
- Access high accuracy vision language model, fine-tune on traffic optimised data and built on NVIDIA Cosmos Reason
- Follow prompt-based instructions for traffic-related operations
- API-first delivery – simple integration via HTTPS
- Fine-tuned models for US and EU markets, with more regions to follow
- Designed to build standalone solutions or integrate with the Milestone product portfolio
- 100% responsibly sourced training data with auditable data lineage, GDPR- and EU AI Act-compliant, used for the fine-tuning of the model
Pricing for the VLMaaS is pay-per-use (based on API calls) – no large upfront investments or custom training costs.
Vision Language Model as a Service
Andrew Burnett, Acting Chief Technology Officer, Milestone Systems, said: “With the Vision Language Model as a Service and Video Summarisation for XProtect, we’re tackling some of the most challenging bottlenecks: video overload and time-consuming manual work. Operators get immediate insight directly within XProtect; builders get API‑first access to production‑ready intelligence without bespoke training or heavy infrastructure."
"Because this model is specialised for real-world traffic video and fine-tuned on responsibly sourced data, customers can trust the results, deploy with confidence, and enhance all existing solutions in place. It’s the fastest, most advanced and impactful path to turning video into actionable outcomes.”
XProtect customers like the cities of Genoa, Italy, and Dubuque, Iowa, US, are excited to use these new capabilities, leading the way in adopting advanced video intelligence solutions to enhance traffic management.
Built on responsible AI, powered by real-world data
The two new offerings are powered by Milestone’s Hafnia VLM, which has been fine-tuned on 75,000 hours of responsibly sourced, real-world video data from either Europe or the US, using NVIDIA Cosmos Curator for data preparation and running either on cloud infrastructure or regional data centres.
Leveraging NVIDIA Cosmos Reason VLM and Milestone’s data for fine-tuning makes it one of the most advanced video AI platforms in the industry.