15 citations · 15 across the 2 of their papers we have counts for
2 papers
cs.CV2023
Learning Instance-Level Representation for Large-Scale Multi-Modal Pretraining in E-commerce
Yang Jin, Yongzhi Li, Zehuan Yuan +1
This paper aims to establish a generic multi-modal foundation model that has the scalable capability to massive downstream applications in E-commerce. Recently, large-scale vision-…
cs.CV2022★ 15 cited
Embracing Consistency: A One-Stage Approach for Spatio-Temporal Video Grounding
Yang Jin, Yongzhi Li, Zehuan Yuan +1
Spatio-Temporal video grounding (STVG) focuses on retrieving the spatio-temporal tube of a specific object depicted by a free-form textual expression. Existing approaches mainly tr…