Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
HERA: Historical Evidence Routing Adapter for Physical Prediction in Latent World Models
Yuanruyi, Yue Cao, Haojia Gao +7
Predictive video models have emerged as promising world models by learning latent visual dynamics from large-scale video. Yet these models remain challenged by physical events unde…
cs.CV2026
PhotoAgent: A Robotic Photographer with Spatial and Aesthetic Understanding
Lirong Che, Zhenfeng Gan, Yanbo Chen +2
Embodied agents for creative tasks like photography must bridge the semantic gap between high-level language commands and geometric control. We introduce PhotoAgent, an agent that…
cs.CV2026
When to Lock Attention: Training-Free KV Control in Video Diffusion
Tianyi Zeng, Jincheng Gao, Tianyi Wang +8
Maintaining background consistency while enhancing foreground quality remains a core challenge in video editing. Injecting full-image information often leads to background artifact…