Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Spatio-Temporal Foundation Models: Vision, Challenges, and Opportunities
Adam Goodge, Wee Siong Ng, Bryan Hooi +1
Foundation models have revolutionized artificial intelligence, setting new benchmarks in performance and enabling transformative capabilities across a wide range of vision and lang…
cs.CV2024
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Lin Xu, Yilin Zhao, Daquan Zhou +3
Vision-language pre-training has significantly elevated performance across a wide range of image-language applications. Yet, the pre-training process for video-related tasks demand…