22 citations · 59 across the 6 of their papers we have counts for
Showing 2025 · cs.CVShow all
2 papers · 2 filters
cs.CV2025
VoCap: Video Object Captioning and Segmentation from Any Prompt
Jasper Uijlings, Xingyi Zhou, Xiuye Gu +5
Understanding objects in videos in terms of fine-grained localization masks and detailed semantic properties is a fundamental task in video understanding. In this paper, we propose…
cs.CV2025★ 1 cited
MALT Diffusion: Memory-Augmented Latent Transformers for Any-Length Video Generation
Sihyun Yu, Meera Hahn, Dan Kondratyuk +6
Diffusion models are successful for synthesizing high-quality videos but are limited to generating short clips (e.g., 2-10 seconds). Synthesizing sustained footage (e.g. over minut…