54 citations · 117 across the 8 of their papers we have counts for
6 papers · 1 filter
Learning Procedure-aware Video Representation from Instructional Videos and Their Narrations
Yiwu Zhong, Licheng Yu, Yang Bai +3
The abundance of instructional videos and their narrations over the Internet offers an exciting avenue for understanding procedural activities. In this work, we propose to learn vi…
Learning and Verification of Task Structure in Instructional Videos
Medhini Narasimhan, Licheng Yu, Sean Bell +2
Given the enormous number of instructional videos available online, learning a diverse array of multi-step task models from videos is an appealing goal. We introduce a new pre-trai…
FAME-ViL: Multi-Tasking Vision-Language Model for Heterogeneous Fashion Tasks
Xiao Han, Xiatian Zhu, Licheng Yu +3
In the fashion domain, there exists a variety of vision-and-language (V+L) tasks, including cross-modal retrieval, text-guided image retrieval, multi-modal classification, and imag…
FashionViL: Fashion-Focused Vision-and-Language Representation Learning
Xiao Han, Licheng Yu, Xiatian Zhu +3
Large-scale Vision-and-Language (V+L) pre-training for representation learning has proven to be effective in boosting various downstream V+L tasks. However, when it comes to the fa…
Detailed Garment Recovery from a Single-View Image
Shan Yang, Tanya Ambert, Zherong Pan +4
Most recent garment capturing techniques rely on acquiring multiple views of clothing, which may not always be readily available, especially in the case of pre-existing photographs…
Modeling Context in Referring Expressions
Licheng Yu, Patrick Poirson, Shan Yang +2
Humans refer to objects in their environments all the time, especially in dialogue with other people. We explore generating and comprehending natural language referring expressions…