Trending: On-device modelsSearch
iHeartGeek
iTECH

NVIDIA's VSS 3.3 cuts the cost of watching every camera

NVIDIA's Video Search and Summarization blueprint adds a skill that composes a working visual AI agent from one prompt in under 30 minutes, plus adaptive sampling that cuts video-language token use by up to 80%.

A figure holding an NVIDIA-branded chip beside three side-by-side video analytics views: a basketball court with player tracking overlays, a grid of warehouse camera feeds and a car park with vehicle detection boxes

NVIDIA has released version 3.3 of its Metropolis blueprint for Video Search and Summarization, or VSS, and the pitch is cost on both ends of a deployment: less work to build a visual AI agent, and less compute to run one. The blueprint connects vision-language models such as NVIDIA Cosmos, language models such as NVIDIA Nemotron, retrieval-augmented generation and Model Context Protocol tools, turning live and recorded video into natural-language search, visual question and answer, verified alerts and automated reporting.

NVIDIA's engineers frame the problem as one of assembly. A production visual AI agent usually spans several workflows at once, so a smart city deployment might need vehicle detection, collision alerting, searchable incident clips, hourly summaries and an operator report, while a warehouse needs people and forklift tracking, near-miss alerts and follow-up questions. Each piece is useful alone, but the value arrives when they run together, and that is where the infrastructure bill, the token bill and the configuration drift all come from.

One prompt, one deployment

The headline addition is a Build Vision Agent skill, published as vss-build-vision-ai, that lets coding agents including Claude Code and Codex compose multi-workflow deployments from a single natural-language request. It starts from one of four validated developer profiles, adds only the services a requested capability actually needs, and converges shared infrastructure such as Kafka, Redis and Elasticsearch onto single instances instead of letting every workflow carry its own copy. NVIDIA's demonstration builds a bottling-line overflow agent: a live, previewable deployment with search, alert verification and shift reporting, assembled in under 30 minutes on a two-GPU RTX PRO 6000 Blackwell host, with the new agent reusing the detector's GPU to run FP8 Cosmos 3 Nano.

The skill also extends a running deployment without rebuilding the whole stack, which NVIDIA frames as the change-cost problem rather than the build-cost one. Adding a capability to a system that is already in production is where proof-of-concept projects usually stall.

Cheaper video at runtime

The second half of the release is Adaptive Efficient Video Sampling, or EVS. Video-language models burn tokens on frames that have not changed, so EVS prunes the visual patches that look the same as the previous frame and batches model work around the moments when something actually happens. On an RTX PRO 6000 Blackwell running Cosmos 3 Super in FP8, NVIDIA reports a 17% cut in alert contextualisation latency and 46% more concurrent real-time VLM streams, and says a 60-minute video was summarised in about half the time using 80% fewer VLM input tokens.

Those numbers come with a caveat NVIDIA states plainly: results vary with scene motion, chunk length and the similarity threshold used. EVS already ships in vLLM and the Cosmos NIM microservices at a fixed pruning rate, while the adaptive version in VSS 3.3 is wired into the real-time VLM microservice and decides per patch and per frame which tokens to keep. The blueprint and its skills are available in NVIDIA's repository, and the company is hosting a live session on 1 October at 9am Pacific to build one from a prompt.

Our opinion

The interesting number in this release is not the 46% or the 80%. It is the under-30-minute build. Visual AI has been technically possible for a while; what kept it out of warehouses and bottling lines was the integration work, and a single prompt that produces a running, previewable deployment with shared infrastructure is a much bigger practical unlock than another model checkpoint. The cynic's read is that this mostly makes NVIDIA's own stack easier to buy, which is true, and also fine, because the pain it removes was real.

The pruning technique deserves more attention than the benchmarks. Sending every frame of a 60-minute video to a model is an expensive way to learn that a warehouse floor was idle for most of it, and Adaptive EVS is a reminder that a lot of the cost of applied AI is not model quality but how much irrelevant data gets pushed through the model in the first place. Expect the technique, or something like it, to show up well beyond video.