2 citations · 3 across the 3 of their papers we have counts for
4 papers
Continual Learning via Sparse Memory Finetuning
Jessy Lin, Luke Zettlemoyer, Gargi Ghosh +4
Modern language models are powerful, but typically static after deployment. A major obstacle to building models that continually learn over time is catastrophic forgetting, where u…
VideoOFA: Two-Stage Pre-Training for Video-to-Text Generation
Xilun Chen, Lili Yu, Wenhan Xiong +3
We propose a new two-stage pre-training framework for video-to-text generation tasks such as video captioning and video question answering: A generative encoder-decoder model is fi…
Text-guided 3D Human Generation from 2D Collections
Tsu-Jui Fu, Wenhan Xiong, Yixin Nie +3
3D human modeling has been widely used for engaging interaction in gaming, film, and animation. The customization of these characters is crucial for creativity and scalability, whi…
Hierarchical Video-Moment Retrieval and Step-Captioning
Abhay Zala, Jaemin Cho, Satwik Kottur +4
There is growing interest in searching for information from large video corpora. Prior works have studied relevant tasks, such as text-based video retrieval, moment retrieval, vide…