6 papers
UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling
An Lanji, Dawei Liu, Jin Li +3
Joint-Embedding Predictive Architectures (JEPAs) have emerged as a principled framework for self-supervised learning of world models in compact latent spaces, yet existing methods…
World2Minecraft: Occupancy-Driven Simulated Scenes Construction
Lechao Zhang, Haoran Xu, Jingyu Gong +3
Embodied intelligence requires high-fidelity simulation environments to support perception and decision-making, yet existing platforms often suffer from data contamination and limi…
GT-Space: Enhancing Heterogeneous Collaborative Perception with Ground Truth Feature Space
Wentao Wang, Haoran Xu, Guang Tan
In autonomous driving, multi-agent collaborative perception enhances sensing capabilities by enabling agents to share perceptual data. A key challenge lies in handling {\em heterog…
Deflickering Vision-Based Occupancy Networks through Lightweight Spatio-Temporal Correlation
Fengcheng Yu, Haoran Xu, Canming Xia +2
Vision-based occupancy networks (VONs) provide an end-to-end solution for reconstructing 3D environments in autonomous driving. However, existing methods often suffer from temporal…
COVR:Collaborative Optimization of VLMs and RL Agent for Visual-Based Control
Canming Xia, Peixi Peng, Guang Tan +4
Visual reinforcement learning (RL) suffers from poor sample efficiency due to high-dimensional observations in complex tasks. While existing works have shown that vision-language m…
Delta-Triplane Transformers as Occupancy World Models
Haoran Xu, Peixi Peng, Guang Tan +3
Occupancy World Models (OWMs) aim to predict future scenes via 3D voxelized representations of the environment to support intelligent motion planning. Existing approaches typically…