activity
20222025
most cited1st Place Solutions for RxR-Habitat Vision-and-Language Navigation Competition (CVPR 2022)

3 citations · 6 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV2025

Pushing the Boundaries of State Space Models for Image and Video Generation

Yicong Hong, Long Mai, Yuan Yao +1

While Transformers have become the dominant architecture for visual generation, linear attention models, such as the state-space models (SSM), are increasingly recognized for their…

cs.CV2024

NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models

Gengze Zhou, Yicong Hong, Zun Wang +2

Capitalizing on the remarkable advancements in Large Language Models (LLMs), there is a burgeoning initiative to harness LLMs for instruction following robotic navigation. Such a t…

cs.CV2024

Augmented Commonsense Knowledge for Remote Object Grounding

Bahram Mohammadi, Yicong Hong, Yuankai Qi +3

The vision-and-language navigation (VLN) task necessitates an agent to perceive the surroundings, follow natural language instructions, and act in photo-realistic unseen environmen…

cs.CV20232 cited

Scaling Data Generation in Vision-and-Language Navigation

Zun Wang, Jialu Li, Yicong Hong +6

Recent research in language-guided visual navigation has demonstrated a significant demand for the diversity of traversable environments and the quantity of supervision for trainin…

cs.CV20231 cited

Learning Navigational Visual Representations with Semantic Map Supervision

Yicong Hong, Yang Zhou, Ruiyi Zhang +4

Being able to perceive the semantics and the spatial structure of the environment is essential for visual navigation of a household robot. However, most existing works only employ…

cs.CV20223 cited

1st Place Solutions for RxR-Habitat Vision-and-Language Navigation Competition (CVPR 2022)

Dong An, Zun Wang, Yangguang Li +5

This report presents the methods of the winning entry of the RxR-Habitat Competition in CVPR 2022. The competition addresses the problem of Vision-and-Language Navigation in Contin…