computer vision

Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks

arXiv:2607.13305

summary

The paper audits video large language model benchmarks by measuring how much their accuracy depends on visual input, introducing the Visual Dependency Gap (VDG) that compares performance on original videos versus black-screen conditions.

Abstract

Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding. We audit this assumption across twenty models spanning 2-78B parameters and ten architecture families. We introduce the Visual Dependency Gap (VDG), the difference in per-question correctness between original-video and black-screen conditions. Paired McNemar tests on MVBench show that accuracy and visual dependency are separable: models differ on original video (p = 0.0003) but not on black screens (p = 0.53). Across models, task-type rankings are stable: Attribute Perception is strongly visual, whereas Temporal Reasoning approaches the language-only baseline. A diagnostic ladder from black screen to single frame, shuffled frames, and original video reveals that frame diversity supplies most of the visual benefit, while temporal order contributes near-zero accuracy across sixteen open-weight models. An ablation from 0.5 to 24 FPS rules out sparse sampling as the cause. H.264 experiments further show that stable aggregate accuracy conceals bidirectional question-level answer flips. The diagnostic also generalizes to four API-accessed models, whose VDG values range from 0.025 to 0.315. These results motivate VDG as a standard audit for whether video benchmarks measure visually grounded capability. Code is available at https://github.com/JaeLee18/accuracy-without-grounding.

Accepted, ACM International Conference on Multimedia 2026 (ACM MM)

Topics & keywords

#video-language models#benchmark auditing#visual dependency#multimodal evaluation#temporal reasoningVisual Dependency GapMVBenchMcNemar testframe diversityblack-screen conditionH.264