2 papers
cs.CV2025
Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
Liping Yuan, Jiawei Wang, Haomiao Sun +2
We introduce Tarsier2, a state-of-the-art large vision-language model (LVLM) designed for generating detailed and accurate video descriptions, while also exhibiting superior genera…
cs.CV2024
Tarsier: Recipes for Training and Evaluating Large Video Description Models
Jiawei Wang, Liping Yuan, Yuchen Zhang +1
Generating fine-grained video descriptions is a fundamental challenge in video understanding. In this work, we introduce Tarsier, a family of large-scale video-language models desi…