Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
FreeSpec: Training-Free Long Video Generation via Singular-Spectrum Reconstruction
Fangda Chen, Shanshan Zhao, Longrong Yang +3
Video diffusion models perform well in short-video synthesis, but their training-free extension to long videos often suffers from content drift, temporal inconsistency, and over-sm…
cs.CV2024
Temporal-Enhanced Multimodal Transformer for Referring Multi-Object Tracking and Segmentation
Changcheng Xiao, Qiong Cao, Yujie Zhong +4
Referring multi-object tracking (RMOT) is an emerging cross-modal task that aims to locate an arbitrary number of target objects and maintain their identities referred by a languag…