8 papers
Compression and Retrieval: Implicit Memory Retrieval for Video World Models
Zhan Peng, Jie Ma, Huiqiang Sun +6
Video world models hold promise for simulating interactive environments, yet maintaining consistent long-term memory across complex camera trajectories remains a critical challenge…
Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
Zhiyu Pan, Yizheng Wu, Jiashen Hua +5
Reasoning has emerged as a key capability of large language models. In linguistic tasks, this capability can be enhanced by self-improving techniques that refine reasoning paths fo…
Semi-Supervised High Dynamic Range Image Reconstructing via Bi-Level Uncertain Area Masking
Wei Jiang, Jiahao Cui, Yizheng Wu +3
Reconstructing high dynamic range (HDR) images from low dynamic range (LDR) bursts plays an essential role in the computational photography. Impressive progress has been achieved b…
SRefiner: Soft-Braid Attention for Multi-Agent Trajectory Refinement
Liwen Xiao, Zhiyu Pan, Zhicheng Wang +2
Accurate prediction of multi-agent future trajectories is crucial for autonomous driving systems to make safe and efficient decisions. Trajectory refinement has emerged as a key st…
Dynamic View Synthesis from Small Camera Motion Videos
Huiqiang Sun, Xingyi Li, Juewen Peng +4
Novel view synthesis for dynamic D scenes poses a significant challenge. Many notable efforts use NeRF-based approaches to address this task and yield impressive results. Howeve…
Deblur-Avatar: Animatable Avatars from Motion-Blurred Monocular Videos
Xianrui Luo, Juewen Peng, Zhongang Cai +4
We introduce a novel framework for modeling high-fidelity, animatable 3D human avatars from motion-blurred monocular video inputs. Motion blur is prevalent in real-world dynamic vi…