TL;DR
Abstract
Large vision-language models (LVLMs) like GPT-4V, Gemini 1.5, Claude, and Qwen 2.5-VL demonstrate impressive performance across various tasks, but can they truly generalize beyond spurious correlations learned during training?
We introduce SpuriVerse, a novel benchmark comprised of 124 distinct types of spurious correlations extracted from real-world datasets, each containing 1 realistic and 10 synthetic VQA samples—1,364 multiple-choice questions in total—sourced from GPT-4o errors on visual question-answering tasks. When tested on 15 open and closed-source models, top performers achieved only 37.1% accuracy. In a separate fine-tuning experiment, Qwen2.5-VL-7B improved from 35.2% to 78.4% accuracy on held-out anchors after training on synthetic examples from other spurious correlation types. This demonstrates generalization to unseen patterns, with a trade-off in performance on non-spurious examples.
Benchmark at a Glance
Key Contributions
- SpuriVerse benchmark: 124 spurious correlation types extracted from real-world datasets, each with realistic and synthetic VQA samples covering a broad range of visual shortcuts.
- Systematic evaluation of 15 LVLMs: Comprehensive assessment of open and closed-source models, revealing that even state-of-the-art systems fall far short on spurious pattern generalization.
- Fine-tuning intervention: Fine-tuning Qwen2.5-VL-7B on diverse synthetic examples improves accuracy on held-out anchors from 35.2% to 78.4%, demonstrating generalization to unseen spurious correlation types.
- Failure mode analysis: Identified when and how models rely on dataset artifacts rather than genuine visual reasoning, providing actionable insights for building more robust VLMs.
Citation
@article{yang2025spuriverse,
title = {Escaping the SpuriVerse: Can Large Vision-Language Models
Generalize Beyond Seen Spurious Correlations?},
author = {Yang, Yiwei and Lee, Chung Peng and Feng, Shangbin and
Zhao, Dora and Wen, Bingbing and Liu, Anthony Z. and
Tsvetkov, Yulia and Howe, Bill},
journal = {arXiv preprint arXiv:2506.18322},
year = {2025}
}