Ambarella, Inc. - Experts & Thought Leaders
Latest Ambarella, Inc. news & announcements
Ambarella, Inc., an edge AI semiconductor company, announced during the ISC West security expo, the continued expansion of its AI system-on-chip (SoC) portfolio with the 5nm CV75S family. These new SoCs provide the industry’s most power- and cost-efficient option for running the latest multi-modal vision-language models (VLMs) and vision-transformer networks. This efficiency makes these cutting-edge AI technologies feasible for a broad range of cost- and power-constrained devices within security cameras for enterprises, smart cities and retail; industrial robotics and access control; and a host of AI-enabled consumer video devices, such as sports and conferencing cameras. Integrating latest technology “With the CV75S family, we are enabling mass-market product designers with the ability to integrate the latest vision-transformer technologies, including VLMs that allow zero shot image classification and multi-modal inferencing for real-time visual analytics without the need for training,” said Chris Day, VP of marketing and business development at Ambarella. “We’re also bringing our advanced AI-based image processing technology to cameras with a wide range of price points, offering significantly greater image quality for a broad spectrum of applications.” Utilising the CV75S This is Ambarella’s first mass-market SoC family to integrate its latest CVflow® 3.0 AI engine A typical example of how the CV75S will be used to run VLMs in enterprise cameras is a natural-language search that is processed within the camera to look for any object or scene among the content it has captured. A multi-modal VLM, such as the contrastive language–image pre-training (CLIP) model, can scour the footage and provide instantaneous results without being trained on that specific object or context. This opens a whole new range of AI capabilities for enterprise cameras, which can now run AI tasks tailored to their installation and user needs without retraining and deploying new AI models for each task. Additional integration This is Ambarella’s first mass-market SoC family to integrate its latest CVflow® 3.0 AI engine, which provides 3x the performance over the prior generation with support for VLMs and vision transformers, as well as advanced AI-based image processing. Additionally, the CV75S integrates the latest generation of Ambarella’s industry-leading image signal processor, 4KP30 H.264/5 video encoding, dual Arm® Cortex-A76 1.6GHz cores, and USB 3.2 connectivity. Ambarella’s Cooper™ Developer Platform To accelerate time to market, the CV75S family is supported by Ambarella’s Cooper™ Developer Platform. This recently introduced platform provides comprehensive hardware and software solutions for creating edge AI systems, including powerful, safe, and secure compute and software capabilities. It consists of industrial-grade hardware tools, collectively called Cooper Metal; along with Cooper Foundry, which provides a multi-layer software stack that supports Ambarella’s entire portfolio of AI SoCs. The CV75S is sampling now and will be demonstrated at Ambarella’s invitation-only exhibition at ISC West in Las Vegas this week.
Ambarella, Inc., an edge AI semiconductor company, announced during the ISC West 2024 security expo, the continued expansion of its AI system-on-chip (SoC) portfolio with the 5nm CV75S family. Power- and cost-efficient option These new SoCs provide the industry’s most power- and cost-efficient option for running the latest multi-modal vision-language models (VLMs) and vision-transformer networks. This efficiency makes these cutting-edge AI technologies feasible for a broad range of cost- and power-constrained devices within security cameras for enterprises, smart cities, and retail; industrial robotics and access control; and a host of AI-enabled consumer video devices, such as sports and conferencing cameras. CV75S family “With the CV75S family, we are enabling mass-market product designers with the ability to integrate the latest vision-transformer technologies, including VLMs that allow zero-shot image classification and multi-modal inferencing for real-time visual analytics without the need for training,” said Chris Day, VP of Marketing and Business Development at Ambarella. He adds, “We’re also bringing our advanced AI-based image processing technology to cameras with a wide range of price points, offering significantly greater image quality for a broad spectrum of applications.” Multi-modal VLM A typical example of how the CV75S will be used to run VLMs in enterprise cameras is a natural-language search A typical example of how the CV75S will be used to run VLMs in enterprise cameras is a natural-language search that is processed within the camera to look for any object or scene among the content it has captured. A multi-modal VLM, such as the contrastive language–image pre-training (CLIP) model, can scour the footage and provide instantaneous results without being trained on that specific object or context. This opens a whole new range of AI capabilities for enterprise cameras, which can run AI tasks tailored to their installation and user needs without retraining and deploying new AI models for each task. CVflow® 3.0 AI engine This is Ambarella’s first mass-market SoC family to integrate its latest CVflow® 3.0 AI engine, which provides 3x the performance over the prior generation with support for VLMs and vision transformers, as well as advanced AI-based image processing. Additionally, the CV75S integrates the latest generation of Ambarella’s industry-leading image signal processor, 4KP30 H.264/5 video encoding, dual Arm® Cortex-A76 1.6GHz cores, and USB 3.2 connectivity. Cooper™ Developer Platform To accelerate time to market, the CV75S family is supported by Ambarella’s Cooper™ Developer Platform. This recently introduced platform provides comprehensive hardware and software solutions for creating edge AI systems, including powerful, safe, and secure computing and software capabilities. It consists of industrial-grade hardware tools, collectively called Cooper Metal; along with Cooper Foundry, which provides a multi-layer software stack that supports Ambarella’s entire portfolio of AI SoCs. The CV75S is sampling and will be demonstrated at Ambarella’s invitation-only exhibition at ISC West event in Las Vegas.
IDS NXT Malibu marks a new class of intelligent industrial cameras that act as edge devices and generate AI overlays in live video streams. For the new camera series, IDS Imaging Development Systems has collaborated with Ambarella, a pioneering developer of visual AI products, making consumer technology available for demanding applications in industrial quality. CVflow® AI vision system It features Ambarella’s CVflow® AI vision system on chip and takes full advantage of the SoC’s advanced image processing and on-camera AI capabilities. Consequently, Image analysis can be performed at high speed (>25fps) and displayed as live overlays in compressed video streams via the RTSP protocol for end devices. SoC’s integrated image signal processor The information captured by the light-sensitive onsemi AR0521 image sensor is processed directly on the camera Due to the SoC’s integrated image signal processor (ISP), the information captured by the light-sensitive onsemi AR0521 image sensor is processed directly on the camera and accelerated by its integrated hardware. The camera also offers helpful automatic features, such as brightness, noise, and colour correction, which significantly improve image quality. Real-time image analysis "With IDS NXT Malibu, we have developed an industrial camera that can analyse images in real-time and incorporate results directly into video streams,” explained Kai Hartmann, Product Innovation Manager at IDS. “The combination of on-camera AI with compression and streaming is a novelty in the industrial setting, opening up new application scenarios for intelligent image processing." Industrial-grade edge AI cameras These on-camera capabilities were made possible through close collaboration between IDS and Ambarella These on-camera capabilities were made possible through close collaboration between IDS and Ambarella, leveraging the companies’ strengths in industrial camera and consumer technology. "We are proud to work with IDS, a leading company in industrial image processing,” said Jerome Gigot, senior director of marketing at Ambarella. “The IDS NXT Malibu represents a new class of industrial-grade edge AI cameras, achieving fast inference times and high image quality via our CVflow AI vision SoC." IDS NXT all-in-one AI system IDS NXT Malibu has entered series production. The camera is part of the IDS NXT all-in-one AI system. Optimally coordinated components from the camera to the AI vision studio accompany the entire workflow. This includes the acquisition of images and their labelling, through to the training of a neural network and its execution on the IDS NXT series of cameras.
Insights & Opinions from thought leaders at Ambarella, Inc.
A security camera installed today has more AI processing power than the systems that guided early autonomous vehicle prototypes. And yet the operator who mounts that camera on a wall will, in all likelihood, never use most of that capability. Industry surveys bear this out: a wide gap persists between the number of security professionals who believe AI can improve outcomes and the much smaller share who have adopted it operationally. The reason has nothing to do with the silicon and everything to do with how the industry has asked people to configure these systems. The problem is not that the industry lacks algorithms. The problem is that physical security has never found a scalable way to personalise systems for each site. The personalisation dilemma hiding in plain sight The problem is that physical security has never found a scalable way to personalise systems for each site A surveillance deployment at an airport, a retail chain, a school campus, and a logistics yard can look strikingly similar in hardware terms. Each installation uses image sensors, edge processors, network connectivity, and a management layer. What changes is what the operator cares about. At a school entrance, the priority might be perimeter approach after hours and controlled access during the day. At a loading dock, the concern is tailgating, vehicle dwell time, and safety incidents near forklifts. At an airport, the operator may need queue-flow analytics one moment, unattended-item detection the next, and then a search for a specific person of interest carrying a particular bag. At a retail store, loss prevention teams want to correlate customer flow patterns with point-of-sale data and identify suspicious behaviour near high-value merchandise. This range of needs forces a reality that the industry has acknowledged in principle but never resolved in practice: the application pool across the market is vast, yet each individual site typically requires only a narrow set of outcomes. Each deployment needs personalisation once, at commissioning, and then again whenever the environment or the risk profile shifts. The app store that never became a market For the better part of a decade, the industry’s most visible answer to the personalisation problem was the “app store” model. The logic was straightforward: curate a marketplace of trained neural network algorithms, let integrators browse a catalogue, and download the right analytic for each job. Queue counting for a passport control hall. License plate recognition for a parking structure. Occupancy monitoring for a conference room. The concept borrowed directly from the consumer smartphone approach. In practice, it never matched physical security’s purchasing and operating rhythm. A phone owner discovers and downloads new apps continuously. A physical security deployment selects one or two analytics functions at installation and rarely revisits them. Another maintenance burden A queue-counting algorithm trained on airport data is excellent at queue counting The economic incentive to maintain, curate, and update a broad catalogue across a fragmented ecosystem of camera OEMs, VMS platforms, and system integrators never materialised when the average buyer drew from only a thin slice of it. And the question of who would operate such a marketplace across that fragmented landscape was never satisfactorily answered. The deeper issue is that distribution was not the hard part. Personalisation was. A queue-counting algorithm trained on airport data is excellent at queue counting. It does not naturally become a general-purpose security tool for whatever the operator needs next. Once a model is trained for a narrow task, adaptation requires another project, another integration cycle, and another maintenance burden. AI-enabled cameras The examples that do exist are instructive. Schiphol Airport in the Netherlands has used trained camera systems for over a decade to measure queue length at passport control and alert staff when additional counters should open. Rome trailed AI-enabled cameras to track pedestrian wait times at crosswalks, measure bus queue length, and monitor parking occupancy to support active transport and reduce vehicle emissions. These are effective, well-regarded deployments. They also illustrate the limitation: each required its own trained model, its own integration effort, and its own maintenance cycle. The queue-counting camera at Schiphol cannot be redeployed to detect an abandoned bag. That is a separate algorithm, a separate procurement, and a separate project. What changes with agentic AI Applied to physical security, this translates into a simpler commissioning experience Agentic AI points to a fundamentally different approach. An agentic system can receive goals expressed in natural language, determine the appropriate actions to fulfil those goals, execute those actions using available tools, and verify the results. Applied to physical security, this translates into a simpler commissioning experience: the operator expresses intent in plain language, and the system configures itself to achieve that intent. Consider the practical implications. An installer commissioning cameras at a retail location could type or speak a set of instructions: “Alert the manager if more than five people are waiting at checkout for longer than two minutes.” A facilities director could ask the system to “Track vehicles that enter the east parking lot after 9 p.m. and flag any that remain for more than 30 minutes.” A school security coordinator might specify: “Notify campus police if anyone approaches the perimeter fence between midnight and 5 a.m.” Appropriate perception capabilities None of these instructions require the operator to select a specific analytic from a catalogue, configure a detection model, or define pixel-level zones in a complex VMS interface. The system interprets the intent, selects the appropriate perception capabilities, configures thresholds and context, and validates behaviour over time. When the operator’s needs change, a new instruction replaces the old one. The camera hardware stays the same. The AI adapts. This is the core of the shift: minimal user input, maximum flexibility, and a security system that personalises itself without requiring the operator to navigate the traditional customise-certify-deploy cycle. Vision language models make it practical A conventional neural network trained for people counting can count people The enabling technology is the vision language model, or VLM. A VLM combines visual encoders with language reasoning, allowing it to interpret images or video in the context of natural language prompts. This is a qualitative leap beyond traditional convolutional neural networks, which classify or detect predefined objects and have no mechanism for open-ended interpretation. A conventional neural network trained for people counting can count people. It cannot distinguish between a crowd of commuters exiting a train station and a crowd assembling in protest. A VLM, by integrating contextual reasoning with visual analysis, can draw inferences that a task-specific model cannot. It can assess behavioural patterns, interpret spatial relationships, and respond to queries about scenes it has never been explicitly trained to analyse. Where a neural network might register two people carrying objects, a VLM could infer whether the scene suggests travellers with luggage or workers transporting equipment, provided the visual context supports that inference. Supporting multimodal input This matters in physical security because operational questions are rarely phrased as taxonomy labels. Operators want to express outcomes. They want to say “show me anything unusual near the loading bay after hours,” and the system should be able to reason about what “unusual” means given the site context. VLMs also support multimodal input. Audio cues such as a raised voice, a scream, an alarm, or breaking glass can contribute to scene interpretation when paired with video. In security applications, where events routinely unfold across both visual and auditory channels, this capability adds a meaningful layer of situational awareness. The edge constraint that forces discipline Large language models in the cloud use hundreds of billions of parameters and consume hundreds of watts None of this works if the architecture assumes data centre conditions. Most surveillance cameras operate under strict power and thermal limits. Power over Ethernet (PoE), the standard delivery mechanism, typically provides between 15 and 30 watts depending on the PoE class, and only a fraction of that budget is available for AI processing after the sensor, ISP, video encoder, and network stack have taken their share. In many installations, the AI workload must fit within a few watts. Large language models in the cloud use hundreds of billions of parameters and consume hundreds of watts. That scale does not translate to a camera mounted on a pole or embedded in a ceiling tile. For agentic AI to work at the edge of a physical security network, the models must be compact, efficient, and designed for the purpose. Neural network acceleration This is where smaller, domain-specific VLMs become essential. Models trained on industry-relevant image and text datasets, combined with techniques such as pruning, quantisation, and parameter-efficient fine-tuning, can deliver meaningful visual reasoning within the compute and memory constraints of an edge processor. The result is a VLM that fits inside a camera’s power budget and still responds to natural language instructions with useful accuracy. Ambarella’s CVflow AI architecture, now in its third generation, was designed for this class of workload. The architecture integrates advanced neural network acceleration with high-resolution image signal processing and video encoding on a single system-on-chip, allowing cameras to run complex AI inference alongside their core imaging functions without exceeding the thermal and power boundaries that define edge deployments. The company's latest addition to its portfolio, the 4-nanometer CV7, runs CNNs and vision language models concurrently across multiple video streams while consuming 20 percent less power than its predecessor. For infrastructure and robotic applications requiring heavier models, the 5-nanometer N1 family supports multimodal LLMs in multi-camera configurations. Distributing intelligence across far edge, near edge, and cloud This tier must respond in milliseconds and operate within a fixed power envelope A workable agentic architecture for physical security distributes intelligence across three tiers, each matched to the processing demands and latency requirements of its role. At the far edge, inside the camera itself, the processor handles real-time perception: object detection, tracking, zone logic, and initial event classification. This tier must respond in milliseconds and operate within a fixed power envelope. At the near edge, on a local gateway or network video recorder, a more capable processor orchestrates across multiple cameras, maintains state, correlates events, retrieves site-specific policies and procedures, and classifies incidents requiring more context than any single camera provides. At the cloud/server tier, available when connectivity permits, the system accesses heavier models for forensic analysis, fleet-wide analytics, model updates, and long-horizon reporting. Periodic cloud access This tiered approach keeps the most time-sensitive decisions local, where latency is lowest and data privacy is strongest. It also means agentic capabilities can scale incrementally. A small installation might run entirely at the far edge with periodic cloud access. A large campus might employ all three tiers, with near-edge orchestration coordinating PTZ patrol patterns across dozens of cameras while the cloud generates shift summaries and updates models based on fleet-wide telemetry. In practice, a security workflow built on this pattern often combines real-time detection at the far edge, behaviour-tree orchestration at the near edge for multi-camera coordination, local retrieval over site playbooks, and conservative safe-mode escalation when system confidence is low. The discipline of deterministic guardrails and structured verification loops is essential in security operations, where unpredictable system behaviour is not acceptable. A hybrid future, with VLMs orchestrating specialist models The transition to agentic AI does not eliminate specialised neural networks The transition to agentic AI does not eliminate specialised neural networks. Purpose-trained models will continue to deliver superior accuracy for well-defined, high-frequency tasks such as license plate recognition, face matching, and fire and smoke detection. In a mature agentic system, the VLM acts as an orchestrator. It handles open-ended perception and natural language interaction while routing to specialised models when a task demands their precision. A PTZ camera at a transportation hub might receive the instruction “monitor the west concourse for unattended items.” The VLM interprets the request, manages the interface, and reasons over broader scene context. Real-time video processing When it identifies a candidate object, it routes to a dedicated abandoned-item classifier optimised for that specific validation step. The VLM orchestrates. The specialist model validates. The operator receives a refined, actionable alert. That hybrid pattern places specific demands on the silicon. The processor must support both traditional CNN inference and generative AI workloads simultaneously while maintaining real-time video processing within the same power envelope. The value of a tightly integrated SoC, one that combines an advanced ISP, a deep learning accelerator, and a video encoder on a single die, is that it eliminates the multi-chip complexity and power overhead that would otherwise make this approach impractical at the edge. Making agentic AI deployable for the ecosystem Ambarella’s Developer Zone, launched at CES 2026, provides a centralised portal of tools Physical security is built on a broad ecosystem of camera OEMs, VMS providers, independent software vendors, module builders, and system integrators. For agentic AI to reach the market at scale, these participants need model-ready tooling, reference workflows, and a practical path from prototype to production. This is where developer ecosystems become part of the story. Ambarella’s Developer Zone, launched at CES 2026, provides a centralised portal of tools, optimised AI models, agentic blueprints, low-code templates, and documentation aimed at accelerating edge AI application development on Ambarella’s SoCs. Common software stack ISVs and integrators can evaluate models, prototype applications, and deploy using a common software stack that spans the company’s CV7 and N1 SoC families through the Cooper development platform. That consistency across the product range reduces per-project engineering cost and accelerates time-to-market for partners building perception and analytics solutions. The point is broader than any single portal: agentic systems require components that have already been tested and optimised for edge deployment, so that integrators can focus on solving their customers' problems rather than rebuilding the AI pipeline from scratch. The ecosystem participants who lead the transition to agentic AI in physical security will be the ones with access to tooling that fits into their existing development and deployment processes. What comes next Physical security has searched for years for a scalable answer to personalisation Physical security has searched for years for a scalable answer to personalisation. The app store model did not provide it. Manual configuration, while functional on a per-site basis, scales poorly across large portfolios of cameras and changing operational requirements. Agentic AI offers a credible path forward because it aligns with how operators actually think. They express outcomes, not model specifications. They want systems that adapt to new requirements without repeated engineering cycles. Traditional neural networks With VLMs as the interface layer, smaller domain-specific models at the far edge, orchestration at the near edge, and disciplined verification loops throughout, personalisation can become a standard part of deployment rather than a custom project. The building blocks are now in place. Power-efficient edge AI processors can run VLMs and traditional neural networks simultaneously. Developer ecosystems are maturing to support rapid prototyping and deployment. Reference architectures for distributing intelligence across far-edge, near-edge, and cloud tiers are solidifying. For an industry that already installs vast numbers of AI-capable cameras each year, the opportunity is to make the intelligence already embedded in those endpoints genuinely usable for the people who rely on them every day.
When it comes to security cameras, the end user always wants more—more resolution, more artificial intelligence (AI), and more sensors. However, the cameras themselves do not change much from generation to generation; that is, they have the same power budgets, form factors and price. To achieve “more,” the systems-on-chips (SoCs) inside the video cameras must pack more features and integrate systems that would have been separate components in the past. For an update on the latest capabilities of SoCs inside video cameras, we turned to Jérôme Gigot, Senior Director of Marketing for AIoT at Ambarella, a manufacturer of SOCs. AIoT refers to the artificial intelligence of things, the combination of AI and IoT. Author's quote “The AI performance on today’s cameras matches what was typically done on a server just a generation ago,” says Gigot. “And, doing AI on-camera provides the threefold benefits of being able to run algorithms on a higher-resolution input before the video is encoded and transferred to a server, with a faster response time, and with complete privacy.” Added features of the new SOC Ambarella expects the first cameras with the SoC to emerge on the market during early part of 2024 Ambarella’s latest System on Chip (SOC) is the CV72S, which provides 6× the AI performance of the previous generation and supports the newer transformer neural networks. Even with its extra features, the CV72S maintains the same power envelope as the previous-generation SoCs. The CV72S is now available, sampling is underway by camera manufacturers, and Ambarella expects the first cameras with the SoC to emerge on the market during the early part of 2024. Examples of the added features of the new SOC include image processing, video encoders, AI engines, de-warpers for fisheye lenses, general compute cores, along with functions such as processing multiple imagers on a single SoC, fusion among different types of sensors, and the list goes on. This article will summarise new AI capabilities based on information provided by Ambarella. AI inside the cameras Gigot says AI is by far the most in-demand feature of new security camera SoCs. Customers want to run the latest neural network architectures; run more of them in parallel to achieve more functions (e.g., identifying pedestrians while simultaneously flagging suspicious behavior); run them at higher resolutions in order to pick out objects that are farther away from the camera. And they want to do it all faster. Most AI tasks can be split between object detection, object recognition, segmentation and higher-level “scene understanding” types of functions, he says. The latest AI engines support transformer network architectures (versus currently used convolutional neural networks). With enough AI horsepower, all objects in a scene can be uniquely identified and classified with a set of attributes, tracked across time and space, and fed into higher-level AI algorithms that can detect and flag anomalies. However, everything depends on which scene is within the camera’s field of view. “It might be an easy task for a camera in an office corridor to track a person passing by every couple of minutes; while a ceiling camera in an airport might be looking at thousands of people, all constantly moving in different directions and carrying a wide variety of bags,” Gigot says. Changing the configuration of video systems Low-level AI number crunching would typically be done on camera (at the source of the data) Even with more computing capability inside the camera, central video servers still have their place in the overall AI deployment, as they can more easily aggregate and understand information across multiple cameras. Additionally, low-level AI number crunching would typically be done on camera (at the source of the data). However, the increasing performance capabilities of transformer neural network AI inside the camera will reduce the need for a central video server over time. Even so, a server could still be used for higher-level decisions and to provide a representation of the world; along with a user interface for the user to make sense of all the data. Overall, AI-enabled security cameras with transformer network-based functionality will greatly reduce the use of central servers in security systems. This trend will contribute to a reduction in the greenhouse gases produced by data centres. These server farms consume a lot of energy, due to their power-hungry GPU and CPU chips, and those server processors also need to be cooled using air conditioning that emits additional greenhouse gases. New capabilities of transformer neural networks New kinds of AI architectures are being deployed inside cameras. Newer SoCs can accommodate the latest transformer neural networks (NNs), which now outperform currently used convolutional NNs for many vision tasks. Transformer neural networks require more AI processing power to run, compared to most convolutional NNs. Transformers are great for Natural Language Processing (NLP) as they have mechanisms to “make sense” of a seemingly random arrangement of words. Those same properties, when applied to video, make transformers very efficient at understanding the world in 3D. Transformer NNs require more AI processing power to run, compared to most convolutional NNs For example, imagine a multi-imager camera where an object needs to be tracked from one camera to the next. Transformer networks are also great at focussing their attention on specific parts of the scene—just as some words are more important than others in a sentence, some parts of a scene might be more significant from a security perspective. “I believe that we are currently just scratching the surface of what can be done with transformer networks in video security applications,” says Gigot. The first use cases are mainly for object detection and recognition. However, research in neural networks is focussing on these new transformer architectures and their applications. Expanded use cases for multi-image and fisheye cameras For multi-image cameras, again, the strategy is “less is more.” For example, if you need to build a multi-imager with four 4K sensors, then, in essence, you need to have four cameras in one. That means you need four imaging pipelines, four encoders, four AI engines, and four sets of CPUs to run the higher-level software and streaming. Of course, for cost, size, and power reasons, it would be extremely inefficient to have four SoCs to do all this processing. Therefore, the latest SoCs for security need to integrate four times the performance of the last generation’s single-imager 4K cameras, in order to process four sensors on a single SoC with all the associated AI algorithms. And they need to do this within a reasonable size and power budget. The challenge is very similar for fisheye cameras, where the SoC needs to be able to accept very high-resolution sensors (i.e., 12MP, 16MP and higher), in order to be able to maintain high resolution after de-warping. Additionally, that same SoC must create all the virtual views needed to make one fisheye camera look like multiple physical cameras, and it has to do all of this while running the AI algorithms on every one of those virtual streams at high resolution. The power of ‘sensor fusion’ Sensor fusion is the ability to process multiple sensor types at the same time and correlate all that information Sensor fusion is the ability to process multiple sensor types at the same time (e.g., visual, radar, thermal and time of flight) and correlate all that information. Performing sensor fusion provides an understanding of the world that is greater than the information that could be obtained from any one sensor type in isolation. In terms of chip design, this means that SoCs must be able to interface with, and natively process, inputs from multiple sensor types. Additionally, they must have the AI and CPU performance required to do either object-level fusion (i.e., matching the different objects identified through the different sensors), or even deep-level fusion. This deep fusion takes the raw data from each sensor and runs AI on that unprocessed data. The result is machine-level insights that are richer than those provided by systems that must first go through an intermediate object representation. In other words, deep fusion eliminates the information loss that comes from preprocessing each individual sensor’s data before fusing it with the data from other sensors, which is what happens in object-level fusion. Better image quality AI can be trained to dramatically improve the quality of images captured by camera sensors in low-light conditions, as well as high dynamic range (HDR) scenes with widely contrasting dark and light areas. Typical image sensors are very noisy at night, and AI algorithms can be trained to perform excellently at removing this noise to provide a clear colour picture—even down to 0.1 lux or below. This is called neural network-based image signal processing, or AISP for short. AI can be trained to perform all these functions with much better results than traditional video methods Achieving high image quality under difficult lighting conditions is always a balance among removing noise, not introducing excessive motion blur, and recovering colours. AI can be trained to perform all these functions with much better results than traditional video processing methods can achieve. A key point for video security is that these types of AI algorithms do not “create” data, they just remove noise and clean up the signal. This process allows AI to provide clearer video, even in challenging lighting conditions. The results are better footage for the humans monitoring video security systems, as well as better input for the AI algorithms analysing those systems, particularly at night and under high dynamic range conditions. A typical example would be a camera that needs to switch to night mode (black and white) when the environmental light falls below a certain lux level. By applying these specially trained AI algorithms, that same camera would be able to stay in colour mode and at full frame rate--even at night. This has many advantages, including the ability to see much farther than a typical external illuminator would normally allow, and reduced power consumption. ‘Straight to cloud’ architecture For the cameras themselves, going to the cloud or to a video management system (VMS) might seem like it doesn’t matter, as this is all just streaming video. However, the reality is more complex; especially for cameras going directly to the cloud. When cameras stream to the cloud, there is usually a mix of local, on-camera storage and streaming, in order to save on bandwidth and cloud storage costs. To accomplish this hybrid approach, multiple video-encoding qualities/resolutions are being produced and sent to different places at the same time; and the camera’s AI algorithms are constantly running to optimise bitrates and orchestrate those different video streams. The ability to support all these different streams, in parallel, and to encode them at the lowest bitrate possible, is usually guided by AI algorithms that are constantly analyzing the video feeds. These are just some of the key components needed to accommodate this “straight to cloud” architecture. Keeping cybersecurity top-of-mind Ambarella’s SoCs always implement the latest security mechanisms, both hardware and software Ambarella’s SoCs always implement the latest security mechanisms, both in hardware and software. They accomplish this through a mix of well-known security features, such as ARM trust zones and encryption algorithms, and also by adding another layer of proprietary mechanisms with things like dynamic random access memory (DRAM) scrambling and key management policies. “We take these measures because cybersecurity is of utmost importance when you design an SoC targeted to go into millions of security cameras across the globe,” says Gigot. ‘Eyes of the world’ – and more brains Cameras are “the eyes of the world,” and visual sensors provide the largest portion of that information, by far, compared to other types of sensors. With AI, most security cameras now have a brain behind those eyes. As such, security cameras have the ability to morph from just a reactive and security-focused apparatus to a global sensing infrastructure that can do everything from regulating the AC in offices based on occupancy, to detecting forest fires before anyone sees them, to following weather and world events. AI is the essential ingredient for the innovation that is bringing all those new applications to life, and hopefully leading to a safer and better world.
Our most popular articles in 2021 provide a good reflection of the state of the industry. Taken together, the Top 10 Articles of 2021, as measured by reader clicks, cover big subjects such as smart cities and cybersecurity. They address new innovations in video surveillance, including systems that are smarter and more connected, and a new generation of computer chips that improve capabilities at the edge. A recurring theme in 2021 is cybersecurity's impact on physical security, embodied by a high-profile hack of 150,000 cameras and an incident at a Florida water plant. There is also an ongoing backlash against facial recognition technology, despite promising technology trends. Cross-agency collaboration Our top articles also touch on subjects that have received less exposure, including use of artificial intelligence (AI) for fraud detection, and the problem of cable theft in South Africa. Here is a review of the Top 10 Articles of 2021, based on reader clicks, including links to the original content: Smart cities have come a long way in the last few decades, but to truly make a smart city safe Safety in Smart Cities: How Video Surveillance Keeps Security Front and Center The main foundations that underpin smart cities are 5G, Artificial Intelligence (AI), and the Internet of Things (IoT) and the Cloud. Each is equally important, and together, these technologies enable city officials to gather and analyse more detailed insights than ever before. For public safety in particular, having IoT and cloud systems in place will be one of the biggest factors to improving the quality of life for citizens. Smart cities have come a long way in the last few decades, but to truly make a smart city safe, real-time situational awareness and cross-agency collaboration are key areas that must be developed as a priority. Fraud detection technology How AI is Revolutionising Fraud Detection Fraud detection technology has advanced rapidly over the years and made it easier for security professionals to detect and prevent fraud. Artificial Intelligence (AI) is revolutionising fraud detection. Banks can use AI software to gain an overview of a customer’s spending habits online. Having this level of insight allows an anomaly detection system to determine whether a transaction is normal or not. Suspicious transactions can be flagged for further investigation and verified by the customer. If the transaction is not fraudulent, then the information can be put into the anomaly detection system to learn more about the customer’s spending behaviour online. For decades, cable theft has caused disruption to infrastructure across South Africa Remote Monitoring Technology: Tackling South Africa’s Cable Theft Problem For decades, cable theft has caused disruption to infrastructure across South Africa, and it’s an issue that permeates the whole supply chain. In November 2020, Nasdaq reported that, “When South Africa shut large parts of its economy and transport network during its COVID-19 lockdown, organised, sometimes armed, gangs moved into its crumbling stations to steal the valuable copper from the lines. Now, more than two months after that lockdown ended, the commuter rail system, relied on by millions of commuters, is barely operational.” Physical security consequences Hack of 150,000 Verkada Cameras: It Could Have Been Worse When 150,000 video surveillance cameras get hacked, it’s big news. The target of the hack was Silicon Valley startup Verkada, which has collected a massive trove of security-camera data from its 150,000 surveillance cameras inside hospitals, companies, police departments, prisons and schools. The data breach was accomplished by an international hacker collective and was first reported by Bloomberg. Water Plant Attack Emphasises Cyber’s Impact on Physical Security At an Oldsmar, Fla., water treatment facility on Feb. 5, an operator watched a computer screen as someone remotely accessed the system monitoring the water supply and increased the amount of sodium hydroxide from 100 parts per million to 11,100 parts per million. The chemical, also known as lye, is used in small concentrations to control acidity in the water. The incident is the latest example of how cybersecurity attacks can translate into real-world, physical security consequences – even deadly ones. Video surveillance technologies Organisations around the globe embraced video surveillance technologies to manage social distancing Video Surveillance is Getting Smarter and More Connected The global pandemic has triggered considerable innovation and change in the video surveillance sector. Last year, organisations around the globe embraced video surveillance technologies to manage social distancing, monitor occupancy levels in internal and external settings, and enhance their return-to-work processes. Forced to reimagine nearly every facet of their operations for a new post-COVID reality, companies were quick to seize on the possibilities offered by today’s next-generation video surveillance systems. The Post-Pandemic Mandate for Entertainment Venues: Digitally Transform Security Guards At sporting venues, a disturbing new trend has hit the headlines — poor fan behaviour. At the same time, security directors are reporting a chronic security guard shortage. Combining surveillance video with AI-based advanced analytics can automatically identify fan disturbances or other operational issues, and notify guards in real time, eliminating the need to have large numbers of guards monitoring video feeds and patrons. The business benefits of digitally transformed guards are compelling. Important emerging technology Why Access Control Is Important In a workspace, access control is particularly crucial in tracking the movement of employees should an incident occur, as well as making the life of your team much easier in allowing them to move between spaces without security personnel and site managers present. It can also reduce the outgoings of a business by reducing the need for security individuals to be hired and paid to remain on site. The city of Baltimore has banned the use of facial recognition systems by residents Baltimore Is the Latest U.S. City to Target Facial Recognition Technology The city of Baltimore has banned the use of facial recognition systems by residents, businesses and the city government (except for police). The criminalisation in a major U.S. city of an important emerging technology in the physical security industry is an extreme example of the continuing backlash against facial recognition throughout the United States. Several localities – from Portland, Oregon, to San Francisco, from Oakland, California, to Boston – have moved to limit use of the technology, and privacy groups have even proposed a national moratorium on use of facial recognition. Powerful artificial intelligence Next Wave of SoCs Will Turbocharge Camera Capabilities at The Edge A new generation of video cameras is poised to boost capabilities dramatically at the edge of the IP network, including more powerful artificial intelligence (AI) and higher resolutions, and paving the way for new applications that would have previously been too expensive or complex. Technologies at the heart of the coming new generation of video cameras are Ambarella’s newest systems on chips (SoCs). Ambarella’s CV5S and CV52S product families are bringing a new level of on-camera AI performance and integration to multi-imager and single-imager IP cameras.
Technology's role in securing banks and financial institutions
DownloadSecurity technologies promote real-time awareness in K-12 schools
DownloadIntegrated systems enable critical and compliant security for transportation
DownloadModernising physical access control
DownloadAccess. Intrusion. One estate.
Download