Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning
Jiahao Wang, An Ping, Yanghai Wang +13
While Omni-modal Large Language Models (OLLMs) have demonstrated impressive capabilities in jointly processing audio and visual streams, their ability to strictly adhere to complex…
cs.CV2025
CauSight: Learning to Supersense for Visual Causal Discovery
Yize Zhang, Meiqi Chen, Sirui Chen +4
Causal thinking enables humans to understand not just what is seen, but why it happens. To replicate this capability in modern AI systems, we introduce the task of visual causal di…