Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
One Ranking, Any Budget: Matryoshka Evidence-to-Context Frame Selection for Long-Video Understanding
Wang Chen, Yu Chen, Xiang Wang +3
Frame selection is essential for applying Large Multimodal Models (LMMs) to long videos due to severe frame redundancy and limited context windows. Since the appropriate frame budg…
cs.CV2026
Credit the Right Box: Marginal Contribution Assignment for Structured Visual Perception
Xinheng Han, Jianfei Wang, Yu Chen +4
Multimodal Large Language Models (MLLMs) are increasingly expected to solve structured perception tasks that require visual recognition, language-to-object binding, object cardinal…
cs.CV2026
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…