3 papers
cs.CV2024
Cross-Modal Few-Shot Learning: a Generative Transfer Learning Framework
Zhengwei Yang, Yuke Li, Qiang Sun +3
Most existing studies on few-shot learning focus on unimodal settings, where models are trained to generalize to unseen data using a limited amount of labeled examples from a singl…
cs.CV2023
Motion Flow Matching for Human Motion Synthesis and Editing
Vincent Tao Hu, Wenzhe Yin, Pingchuan Ma +7
Human motion synthesis is a fundamental task in computer animation. Recent methods based on diffusion models or GPT structure demonstrate commendable performance but exhibit drawba…
cs.CL2023
Semi-supervised multimodal coreference resolution in image narrations
Arushi Goel, Basura Fernando, Frank Keller +1
In this paper, we study multimodal coreference resolution, specifically where a longer descriptive text, i.e., a narration is paired with an image. This poses significant challenge…