6 papers
Plan in Sandbox, Navigate in Open Worlds: Learning Physics-Grounded Abstracted Experience for Embodied Navigation
Zhixuan Shen, Jiawei Du, Ziyu Guo +5
Vision-Language Models (VLMs) have demonstrated exceptional general reasoning capabilities. However, their performance in embodied navigation remains hindered by a scarcity of alig…
Plug-and-Play Label Map Diffusion for Universal Goal-Oriented Navigation
Zhixuan Shen, Yijie Zeng, Shengxiang Luo +2
In embodied vision, Goal-Oriented Navigation (GON) requires robots to locate a specific goal within an unexplored environment. The primary challenge of GON arises from the need to…
Unlocking air traffic flow prediction through microscopic aircraft-state modeling
Bin Wang, Anqi Liu, Jiangtao Zhao +8
Short-term air traffic flow prediction in terminal airspace is essential for proactive air traffic management. Existing approaches predominantly model traffic flow as aggregated ti…
Harnessing Meta-Learning for Controllable Full-Frame Video Stabilization
Muhammad Kashif Ali, Eun Woo Im, Dongjin Kim +4
Video stabilization remains a fundamental problem in computer vision, particularly pixel-level synthesis solutions for video stabilization, which synthesize full-frame outputs, add…
Enhancing Multi-Robot Semantic Navigation Through Multimodal Chain-of-Thought Score Collaboration
Zhixuan Shen, Haonan Luo, Kexun Chen +2
Understanding how humans cooperatively utilize semantic knowledge to explore unfamiliar environments and decide on navigation directions is critical for house service multi-robot s…
Adversarial Training with OCR Modality Perturbation for Scene-Text Visual Question Answering
Zhixuan Shen, Haonan Luo, Sijia Li +1
Scene-Text Visual Question Answering (ST-VQA) aims to understand scene text in images and answer questions related to the text content. Most existing methods heavily rely on the ac…