2 citations · 2 across the 3 of their papers we have counts for
1 paper
Liping Yuan, Jiawei Wang, Haomiao Sun +2
We introduce Tarsier2, a state-of-the-art large vision-language model (LVLM) designed for generating detailed and accurate video descriptions, while also exhibiting superior genera…