5 citations · 7 across the 2 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Haokun Lin, Teng Wang, Yixiao Ge +6
Pioneering token-based works such as Chameleon and Emu3 have established a foundation for multimodal unification but face challenges of high training computational overhead and lim…
cs.CV2022★ 2 cited
Semantic-Aware Pretraining for Dense Video Captioning
Teng Wang, Zhu Liu, Feng Zheng +3
This report describes the details of our approach for the event dense-captioning task in ActivityNet Challenge 2021. We present a semantic-aware pretraining method for dense video…
cs.CV2020★ 5 cited
Dense-Captioning Events in Videos: SYSU Submission to ActivityNet Challenge 2020
Teng Wang, Huicheng Zheng, Mingjing Yu
This technical report presents a brief description of our submission to the dense video captioning task of ActivityNet Challenge 2020. Our approach follows a two-stage pipeline: fi…