Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Long Video Understanding with Learnable Retrieval in Video-Language Models
Jiaqi Xu, Cuiling Lan, Wenxuan Xie +2
The remarkable natural language understanding, reasoning, and generation capabilities of large language models (LLMs) have made them attractive for application to video understandi…
cs.CV2024
GSemSplat: Generalizable Semantic 3D Gaussian Splatting from Uncalibrated Image Pairs
Xingrui Wang, Cuiling Lan, Hanxin Zhu +2
Modeling and understanding the 3D world is crucial for various applications, from augmented reality to robotic navigation. Recent advancements based on 3D Gaussian Splatting have i…
cs.CV2024
TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video Generation
Xingrui Wang, Xin Li, Yaosi Hu +4
Text-driven Image to Video Generation (TI2V) aims to generate controllable video given the first frame and corresponding textual description. The primary challenges of this task li…