activity
20212024
most citedAdversarial Reinforced Instruction Attacker for Robust Vision-Language Navigation

19 citations · 37 across the 12 of their papers we have counts for

collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV20245 cited

Towards Deviation-Robust Agent Navigation via Perturbation-Aware Contrastive Learning

Bingqian Lin, Yanxin Long, Yi Zhu +4

Vision-and-language navigation (VLN) asks an agent to follow a given language instruction to navigate through a real 3D environment. Despite significant advances, conventional VLN…

cs.CV2024

DNA Family: Boosting Weight-Sharing NAS with Block-Wise Supervisions

Guangrun Wang, Changlin Li, Liuchun Yuan +5

Neural Architecture Search (NAS), aiming at automatically designing neural architectures by machines, has been considered a key step toward automatic machine learning. One notable…

cs.CV2024

AlignMiF: Geometry-Aligned Multimodal Implicit Field for LiDAR-Camera Joint Synthesis

Tao Tang, Guangrun Wang, Yixing Lao +5

Neural implicit fields have been a de facto standard in novel view synthesis. Recently, there exist some methods exploring fusing multiple modalities within a single field, aiming…

cs.CV20231 cited

DiffCloth: Diffusion Based Garment Synthesis and Manipulation via Structural Cross-modal Semantic Alignment

Xujie Zhang, Binbin Yang, Michael C. Kampffmeyer +6

Cross-modal garment synthesis and manipulation will significantly benefit the way fashion designers generate garments and modify their designs via flexible linguistic interfaces.Cu…

cs.CV2023

LAW-Diffusion: Complex Scene Generation by Diffusion with Layouts

Binbin Yang, Yi Luo, Ziliang Chen +3

Thanks to the rapid development of diffusion models, unprecedented progress has been witnessed in image synthesis. Prior works mostly rely on pre-trained linguistic models, but a t…

cs.CV2023

Actional Atomic-Concept Learning for Demystifying Vision-Language Navigation

Bingqian Lin, Yi Zhu, Xiaodan Liang +2

Vision-Language Navigation (VLN) is a challenging task which requires an agent to align complex visual observations to language instructions to reach the goal position. Most existi…