7 citations · 15 across the 8 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2023★ 7 cited
Learning Grounded Vision-Language Representation for Versatile Understanding in Untrimmed Videos
Teng Wang, Jinrui Zhang, Feng Zheng +3
Joint video-language learning has received increasing attention in recent years. However, existing works mainly focus on single or multiple trimmed video clips (events), which make…
cs.CV2022
Exploiting Context Information for Generic Event Boundary Captioning
Jinrui Zhang, Teng Wang, Feng Zheng +2
Generic Event Boundary Captioning (GEBC) aims to generate three sentences describing the status change for a given time boundary. Previous methods only process the information of a…