2 papers
cs.CV2026
UniFusion: Sparse-View 4D Reconstruction via Unified Spatio-temporal Depth Alignment
Yongzhe Lyu, Shaofei Wang, Yixin Chen +1
In this paper, we address the challenging problem of 4D reconstruction from sparse-view videos. This setup usually relies on monocular depth estimation to provide priors for the re…
cs.CV2026
Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing
Gengtian Shi, Jinze Yu, Chenhao Wu +5
Video-text temporal localization requires precise alignment between natural language queries and corresponding video segments, a fundamental challenge in multimodal understanding.…