activity
20242026
collaborators

6 papers

cs.CV2026

UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling

An Lanji, Dawei Liu, Jin Li +3

Joint-Embedding Predictive Architectures (JEPAs) have emerged as a principled framework for self-supervised learning of world models in compact latent spaces, yet existing methods…

cs.CV2026

World2Minecraft: Occupancy-Driven Simulated Scenes Construction

Lechao Zhang, Haoran Xu, Jingyu Gong +3

Embodied intelligence requires high-fidelity simulation environments to support perception and decision-making, yet existing platforms often suffer from data contamination and limi…

cs.LG2026

GT-Space: Enhancing Heterogeneous Collaborative Perception with Ground Truth Feature Space

Wentao Wang, Haoran Xu, Guang Tan

In autonomous driving, multi-agent collaborative perception enhances sensing capabilities by enabling agents to share perceptual data. A key challenge lies in handling {\em heterog…

cs.CV2026

Deflickering Vision-Based Occupancy Networks through Lightweight Spatio-Temporal Correlation

Fengcheng Yu, Haoran Xu, Canming Xia +2

Vision-based occupancy networks (VONs) provide an end-to-end solution for reconstructing 3D environments in autonomous driving. However, existing methods often suffer from temporal…

cs.CV2026

COVR:Collaborative Optimization of VLMs and RL Agent for Visual-Based Control

Canming Xia, Peixi Peng, Guang Tan +4

Visual reinforcement learning (RL) suffers from poor sample efficiency due to high-dimensional observations in complex tasks. While existing works have shown that vision-language m…

cs.RO2025

Delta-Triplane Transformers as Occupancy World Models

Haoran Xu, Peixi Peng, Guang Tan +3

Occupancy World Models (OWMs) aim to predict future scenes via 3D voxelized representations of the environment to support intelligent motion planning. Existing approaches typically…