activity
20202024
most citedHOP: History-and-Order Aware Pre-training for Vision-and-Language Navigation

8 citations · 8 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV2024

MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation

Mingzhen Sun, Weining Wang, Yanyuan Qiao +5

Sounding Video Generation (SVG) is an audio-video joint generation task challenged by high-dimensional signal spaces, distinct data formats, and different patterns of content infor…

cs.RO2024

Effective Tuning Strategies for Generalist Robot Manipulation Policies

Wenbo Zhang, Yang Li, Yanyuan Qiao +5

Generalist robot manipulation policies (GMPs) have the potential to generalize across a wide range of tasks, devices, and environments. However, existing policies continue to strug…

cs.CV2024

MiniVLN: Efficient Vision-and-Language Navigation by Progressive Knowledge Distillation

Junyou Zhu, Yanyuan Qiao, Siqi Zhang +3

In recent years, Embodied Artificial Intelligence (Embodied AI) has advanced rapidly, yet the increasing size of models conflicts with the limited computational capabilities of Emb…

cs.CV20228 cited

HOP: History-and-Order Aware Pre-training for Vision-and-Language Navigation

Yanyuan Qiao, Yuankai Qi, Yicong Hong +3

Pre-training has been adopted in a few of recent works for Vision-and-Language Navigation (VLN). However, previous pre-training methods for VLN either lack the ability to predict f…

cs.CV2020

Referring Expression Comprehension: A Survey of Methods and Datasets

Yanyuan Qiao, Chaorui Deng, Qi Wu

Referring expression comprehension (REC) aims to localize a target object in an image described by a referring expression phrased in natural language. Different from the object det…