activity
20242026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV20252 cited

Unseen from Seen: Rewriting Observation-Instruction Using Foundation Models for Augmenting Vision-Language Navigation

Ziming Wei, Bingqian Lin, Yunshuang Nie +4

Data scarcity is a long-standing challenge in the Vision-Language Navigation (VLN) field, which extremely hinders the generalization of agents to unseen environments. Previous work…

cs.CV2025

Structured Preference Optimization for Vision-Language Long-Horizon Task Planning

Xiwen Liang, Min Lin, Weiqi Ruan +6

Existing methods for vision-language task planning excel in short-horizon tasks but often fall short in complex, long-horizon planning within dynamic environments. These challenges…

cs.CV2025

NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning

Bingqian Lin, Yunshuang Nie, Ziming Wei +6

Vision-and-Language Navigation (VLN), as a crucial research problem of Embodied AI, requires an embodied agent to navigate through complex 3D environments following natural languag…

cs.CV2025

EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Kai Chen, Yunhao Gou, Runhui Huang +28

GPT-4o, an omni-modal model that enables vocal conversations with diverse emotions and tones, marks a milestone for omni-modal foundation models. However, empowering Large Language…

cs.CV2024

Correctable Landmark Discovery via Large Models for Vision-Language Navigation

Bingqian Lin, Yunshuang Nie, Ziming Wei +5

Vision-Language Navigation (VLN) requires the agent to follow language instructions to reach a target position. A key factor for successful navigation is to align the landmarks imp…

cs.CV2024

Language-Driven Visual Consensus for Zero-Shot Semantic Segmentation

Zicheng Zhang, Tong Zhang, Yi Zhu +4

The pre-trained vision-language model, exemplified by CLIP, advances zero-shot semantic segmentation by aligning visual features with class embeddings through a transformer decoder…