13 papers
NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks
Zhihao Luo, Wentao Yan, Jingyu Gong +5
Recent advances in Graphical User Interface (GUI) and embodied navigation have driven progress, yet these domains have largely evolved in isolation, with disparate datasets and tra…
Vision-language models lag human performance on physical dynamics and intent reasoning
Tianjun Gu, Jingyu Gong, Zhizhong Zhang +4
Spatial intelligence is central to embodied cognition, yet contemporary AI systems still struggle to reason about physical interactions in open-world human environments. Despite st…
Continual-NExT: A Unified Comprehension And Generation Continual Learning Framework
Jingyang Qiao, Zhizhong Zhang, Xin Tan +3
Dual-to-Dual MLLMs refer to Multimodal Large Language Models, which can enable unified multimodal comprehension and generation through text and image modalities. Although exhibitin…
Diffusion Implicit Policy for Unpaired Scene-aware Motion Synthesis
Jingyu Gong, Chong Zhang, Fengqi Liu +5
Scene-aware motion synthesis has been widely researched recently due to its numerous applications. Prevailing methods rely heavily on paired motion-scene data, while it is difficul…
PFDepth: Heterogeneous Pinhole-Fisheye Joint Depth Estimation via Distortion-aware Gaussian-Splatted Volumetric Fusion
Zhiwei Zhang, Ruikai Xu, Weijian Zhang +5
In this paper, we present the first pinhole-fisheye framework for heterogeneous multi-view depth estimation, PFDepth. Our key insight is to exploit the complementary characteristic…
DORAEMON: Decentralized Ontology-aware Reliable Agent with Enhanced Memory Oriented Navigation
Tianjun Gu, Linfeng Li, Xuhong Wang +6
Adaptive navigation in unfamiliar environments is crucial for household service robots but remains challenging due to the need for both low-level path planning and high-level scene…