75 citations · 201 across the 19 of their papers we have counts for
6 papers · 2 filters
Being Comes from Not-being: Open-vocabulary Text-to-Motion Generation with Wordless Training
Junfan Lin, Jianlong Chang, Lingbo Liu +4
Text-to-motion generation is an emerging and challenging problem, which aims to synthesize motion with the same semantics as the input text. However, due to the lack of diverse lab…
Retrospectives on the Embodied AI Workshop
Matt Deitke, Dhruv Batra, Yonatan Bisk +36
We present a retrospective on the state of Embodied AI research. Our analysis focuses on 13 challenges presented at the Embodied AI Workshop at CVPR. These challenges are grouped i…
Towards a Unified View on Visual Parameter-Efficient Transfer Learning
Bruce X. B. Yu, Jianlong Chang, Lingbo Liu +2
Parameter efficient transfer learning (PETL) aims at making good use of the representation knowledge in the pre-trained large models by fine-tuning a small number of parameters. Re…
Taking an Emotional Look at Video Paragraph Captioning
Qinyu Li, Tengpeng Li, Hanli Wang +1
Translating visual data into natural language is essential for machines to understand the world and interact with humans. In this work, a comprehensive study is conducted on video…
Knowledge-enriched Attention Network with Group-wise Semantic for Visual Storytelling
Tengpeng Li, Hanli Wang, Bin He +1
As a technically challenging topic, visual storytelling aims at generating an imaginary and coherent story with narrative multi-sentences from a group of relevant images. Existing…
Visual Acoustic Matching
Changan Chen, Ruohan Gao, Paul Calamia +1
We introduce the visual acoustic matching task, in which an audio clip is transformed to sound like it was recorded in a target environment. Given an image of the target environmen…