6 papers
SkillNav: Score-Level Skill Intervention for Zero-Shot Object Goal Navigation
Ruijie Sang, Yiqun Duan, Pinhan Fu +3
Vision-Language Model (VLM) agents have advanced zero-shot object-goal navigation, yet single-frame reasoning leaves them without the cross-step behavioral awareness an embodied na…
Humanoid-OmniOcc: Stereo-Based Full-View Occupancy Dataset for Embodied AI
Xianda Guo, Bohao Zhang, Chenwei Huang +6
Occupancy prediction at voxel-level granularity is essential for safe robotic navigation and interaction in complex environments. Existing occupancy datasets, however, are predomin…
ROVR-Open-Dataset: A Large-Scale Depth Dataset for Autonomous Driving
Xianda Guo, Ruijun Zhang, Yiqun Duan +9
Depth estimation is a fundamental component of spatial perception for autonomous driving and other unmanned systems operating in open urban environments. Existing depth datasets su…
Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers
Xingyue Huang, Rishabh, Gregor Franke +43
Recent advances in Large Language Models (LLMs) have shown that their reasoning capabilities can be significantly improved through Reinforcement Learning with Verifiable Reward (RL…
Scaling Prompt Instructed Zero Shot Composed Image Retrieval with Image-Only Data
Yiqun Duan, Sameera Ramasinghe, Stephen Gould +1
Composed Image Retrieval (CIR) is the task of retrieving images matching a reference image augmented with a text, where the text describes changes to the reference image in natural…
Agent-Centric Personalized Multiple Clustering with Multi-Modal LLMs
Ziye Chen, Yiqun Duan, Riheng Zhu +2
Personalized multiple clustering aims to generate diverse partitions of a dataset based on different user-specific aspects, rather than a single clustering. It has recently drawn r…