Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Long Video Understanding with Learnable Retrieval in Video-Language Models
Jiaqi Xu, Cuiling Lan, Wenxuan Xie +2
The remarkable natural language understanding, reasoning, and generation capabilities of large language models (LLMs) have made them attractive for application to video understandi…
cs.CV2025
LTM3D: Bridging Token Spaces for Conditional 3D Generation with Auto-Regressive Diffusion Framework
Xin Kang, Zihan Zheng, Lei Chu +5
We present LTM3D, a Latent Token space Modeling framework for conditional 3D shape generation that integrates the strengths of diffusion and auto-regressive (AR) models. While diff…
cs.CV2024
Slot-VLM: SlowFast Slots for Video-Language Modeling
Jiaqi Xu, Cuiling Lan, Wenxuan Xie +2
Video-Language Models (VLMs), powered by the advancements in Large Language Models (LLMs), are charting new frontiers in video understanding. A pivotal challenge is the development…