9 citations · 15 across the 2 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2023★ 9 cited
Joint Visual Grounding and Tracking with Natural Language Specification
Li Zhou, Zikun Zhou, Kaige Mao +1
Tracking by natural language specification aims to locate the referred target in a sequence based on the natural language description. Existing algorithms solve this issue in two s…
cs.CV2021★ 6 cited
Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational Autoencoders
Jing Li, Di Kang, Wenjie Pei +4
Generating conversational gestures from speech audio is challenging due to the inherent one-to-many mapping between audio and body motions. Conventional CNNs/RNNs assume one-to-one…