12 citations · 17 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 3 cited
VideoLLM-online: Online Video Large Language Model for Streaming Video
Joya Chen, Zhaoyang Lv, Shiwei Wu +7
Recent Large Language Models have been enhanced with vision capabilities, enabling them to comprehend images, videos, and interleaved vision-language content. However, the learning…
cs.CV2024★ 2 cited
Towards A Better Metric for Text-to-Video Generation
Jay Zhangjie Wu, Guian Fang, Haoning Wu +11
Generative models have demonstrated remarkable capability in synthesizing high-quality text, images, and videos. For video generation, contemporary text-to-video models exhibit imp…
cs.CV2023★ 12 cited
Sounding Video Generator: A Unified Framework for Text-guided Sounding Video Generation
Jiawei Liu, Weining Wang, Sihan Chen +2
As a combination of visual and audio signals, video is inherently multi-modal. However, existing video generation methods are primarily intended for the synthesis of visual frames,…