- Updated: July 4, 2026
- 2 min read
Robusto‑2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City
Robusto‑2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City
Self‑driving cars are rapidly expanding across the globe, and modern autonomous systems increasingly rely on multimodal models such as Vision‑Language Models (VLMs) to interpret complex driving environments. The Robusto‑2 study, presented on arXiv, provides the first large‑scale, factorial analysis of how both human drivers and VLMs perform on visual question‑answering (VQA) tasks in two of the most challenging urban settings: Lima, Peru and New York City, USA.

Why Lima and New York City?
Both cities represent extreme edge‑case scenarios for autonomous driving. They feature dense traffic, unpredictable pedestrian behavior, and a lack of existing self‑driving car deployments. By selecting these locations, the authors ensure that the dataset captures out‑of‑distribution (OOD) conditions that are critical for evaluating the robustness of VLM‑based driving agents.
Key Findings
- Human‑VLM Divergence: Responses from humans and VLMs differ significantly, especially depending on the question category (Factual, Ratings, Counterfactual, Reasoning).
- Geographic Consistency: Human drivers answered similarly regardless of whether they were from Lima or NYC, suggesting that human cognition is less sensitive to geographic context in this VQA framework.
- Limited Geographic Effect on VLMs: Contrary to expectations, VLM performance did not vary strongly with geography, likely because the questions target highly OOD visual content.
Dataset Availability
The complete dataset, including dashcam footage and annotated Q&A pairs, is publicly available on Hugging Face. Researchers can download the data to explore new VLM architectures, improve domain adaptation techniques, or develop safety‑critical validation pipelines.
Implications for Autonomous Driving
Robusto‑2 highlights the importance of:
- Designing VLMs that can handle diverse, real‑world visual inputs beyond curated datasets.
- Incorporating human‑in‑the‑loop evaluation to benchmark AI systems against expert drivers.
- Building geographic‑agnostic perception modules that maintain performance across varied urban landscapes.
For more insights on how Ubos Tech integrates cutting‑edge VLM research into autonomous driving solutions, visit our blog or learn about our mission.
Author: Ubos Tech Team
Andrii Bidochko
CTO UBOS
Andrii Bidochko is an AI entrepreneur and researcher focused on AI agents, reinforcement learning, and autonomous systems. He writes about the technologies shaping the future of machine intelligence, from frontier models and agent architectures to real-world AI applications.