The rapid advancement of multimodal AI is reshaping how intelligent systems interact with the physical world. Vision language models (VLMs) are no longer limited to cloud-based assistants—they are increasingly being deployed in warehouses, manufacturing facilities, traffic monitoring systems, robotics, and other physical AI applications where decisions must be made in real time.
As these deployments grow, so do the engineering challenges. Success is no longer measured solely by reasoning capability. Models must deliver that capability while operating within strict constraints on latency, memory, power, and compute. As the field evolves, so must the benchmarks used to evaluate it.
To keep pace with this shift, VANTAGE-Bench now includes NVIDIA Cosmos3 Edge alongside three additional compact reasoning models, expanding coverage of the latest generation of deployment-focused VLMs. By continuously incorporating emerging model families into a consistent operational evaluation framework, VANTAGE-Bench provides researchers and practitioners with timely, reproducible evaluations that reflect the current state of multimodal AI.
Newly Added Models
- NVIDIA Cosmos3-Edge
- NVIDIA Cosmos-Reason2-2B
- Gemma-4-E2B-it
- Qwen3-VL-2B-Instruct

Figure 1. Public leaderboard highlighting the newly added models.
Why Edge Models Need Their Own Evaluation
Edge models aren’t simply smaller versions of frontier-scale systems—they’re designed around a fundamentally different engineering objective.
Rather than maximizing reasoning capability with virtually unlimited compute, edge models must maximize reasoning within a fixed deployment budget. Latency, memory footprint, power consumption, and hardware availability all become part of the optimization problem. Whether deployed on warehouse robots, manufacturing inspection systems, intelligent traffic cameras, or autonomous platforms, these models often need to process information where it is generated instead of relying on cloud-scale infrastructure.
That changes what a benchmark should measure. It’s no longer just how well a model reasons, but how effectively it reasons under realistic operational constraints. Evaluating compact reasoning models alongside larger systems provides a more complete understanding of the trade-offs between capability, efficiency, and deployability—an increasingly important consideration as physical AI moves from research to production.
This update represents more than the addition of four reasoning models—it reflects how rapidly the multimodal AI ecosystem is evolving and the need for benchmarks to evolve alongside it.
Among the newly evaluated models, Cosmos3 Edge as a reasoning vision language model (VLM) achieved the highest overall performance within the evaluated ~2B parameter class, establishing a strong baseline for this emerging generation of compact vision-language models. More importantly, evaluating these models under the same operational tasks and methodology enables meaningful comparisons within similar deployment budgets, helping researchers better understand the trade-offs between model capability and deployment efficiency. Full results for all four models are available on the VANTAGE-Bench leaderboard.
At Clemson, we view VANTAGE-Bench as research infrastructure rather than a static benchmark. Maintaining an operational benchmark is an ongoing effort—one that requires continuously integrating new model families, architectures, and deployment paradigms as the field advances. As Physical AI continues to evolve, VANTAGE-Bench will continue evolving with it, providing the research community with a consistent, reproducible framework for evaluating the next generation of vision-language models.