benchmark auditing 1multimodal evaluation 1temporal reasoning 1video-language models 1visual dependency 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CV2026
Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks
Jae Joong Lee
The paper audits video large language model benchmarks by measuring how much their accuracy depends on visual input, introducing the Visual Dependency Gap (VDG) that compares perfo…
cs.CV2026
Language-Guided Invariance Probing of Vision-Language Models
Jae Joong Lee
Recent vision-language models (VLMs) such as CLIP, OpenCLIP, EVA02-CLIP and SigLIP achieve strong zero-shot performance, but it is unclear how reliably they respond to controlled l…