4 papers
The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping
Sarvesh Baskar, Zikui Cai, Shayan Shabihi +5
Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, making failure modes hard to isolate. While e…
Compositional Adversarial Training for Robust Visual Watermarking
Anirudh Satheesh, Michael-Andrei Panaitescu-Liess, Andrew Xu +4
Robust watermarking is typically trained with random post-processing augmentation, but random sampling under-covers the combinatorial space of realistic attack pipelines and rarely…
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
Ang Li, Charles Wang, Deqing Fu +9
Humans often use visual aids, for example diagrams or sketches, when solving complex problems. Training multimodal models to do the same, known as Visual Chain of Thought (Visual C…
Zero-Shot Vision Encoder Grafting via LLM Surrogates
Kaiyu Yue, Vasu Singla, Menglin Jia +6
Vision language models (VLMs) typically pair a modestly sized vision encoder with a large language model (LLM), e.g., Llama-70B, making the decoder the primary computational burden…