Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning
Jiahao Wang, An Ping, Yanghai Wang +13
While Omni-modal Large Language Models (OLLMs) have demonstrated impressive capabilities in jointly processing audio and visual streams, their ability to strictly adhere to complex…
cs.CV2025
PAN: A World Model for General, Interactable, and Long-Horizon World Simulation
PAN Team, Jiannan Xiang, Yi Gu +31
A world model enables an intelligent agent to imagine, predict, and reason about how the world evolves in response to its actions, and accordingly to plan and strategize. While rec…