collaborators

6 papers

cs.CV2026

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models

Haodong Yan, Junfeng Li, Junjie He +12

Mainstream World-Action Models (WAMs) adapt pretrained video generation models (VGMs) for robot control, transferring their learned dynamics prior for action prediction. These VGMs…

cs.CV2026

InfiniVerse: Occupancy Guided Unbounded Scene Generation for Autonomous Driving

Xiaoyu Ye, Leheng Li, Xinyu Ji +8

Generating realistic, controllable, and temporally coherent urban environments is a critical yet unresolved challenge in the autonomous driving community. In this paper, we introdu…

cs.CV2026

DualCoT-VLA: Visual-Linguistic Chain of Thought via Parallel Reasoning for Vision-Language-Action Models

Zhide Zhong, Junfeng Li, Junjie He +10

Vision-Language-Action (VLA) models map visual observations and language instructions directly to robotic actions. While effective for simple tasks, standard VLA models often strug…

cs.CV2026

S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight

Haodong Yan, Zhide Zhong, Jiaguan Zhu +10

Video action models (VAMs) have emerged as a promising paradigm for robot learning, owing to their powerful visual foresight for complex manipulation tasks. However, current VAMs,…

q-bio.NC2026

Inferring brain plasticity rule under long-term stimulation with structured recurrent dynamics

Zhichao Liang, Jingzhe Lin, Xinyi Li +2

Understanding how long-term stimulation reshapes neural circuits requires uncovering the rules of brain plasticity. While short-term synaptic modifications have been extensively ch…

cs.CV2025

SATMapTR: Satellite Image Enhanced Online HD Map Construction

Bingyuan Huang, Guanyi Zhao, Qian Xu +3

High-definition (HD) maps are evolving from pre-annotated to real-time construction to better support autonomous driving in diverse scenarios. However, this process is hindered by…