activity
20242026
collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2026

WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving

Xinlin Wang, Yujiao Xiang, Yuheng Zhou +11

Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent feature prediction. However, V-JEPA…

cs.CV2026

XDG: Accelerated Visual Disambiguation

Gonglin Chen, Ben Southall, Hanyuan Xiao +8

Visual aliasing, also known as the doppelganger problem, remains a key challenge for structure-from-motion (SfM): visually similar but physically distinct surfaces can produce inco…

cs.CV2026

ChainFlow-VLA: Causal Flow Planning with Vision-Language Models

Xiyang Wang, Xinlin Wang, Tingguang Zhou +7

Current end-to-end autonomous driving systems are fundamentally limited by a mismatch between temporal causal reasoning and global trajectory consistency. Autoregressive (AR) model…

cs.CV2026

DCARL: A Divide-and-Conquer Framework for Autoregressive Long-Trajectory Video Generation

Junyi Ouyang, Wenbin Teng, Gonglin Chen +2

Long-trajectory video generation is a crucial yet challenging task for world modeling primarily due to the limited scalability of existing video diffusion models (VDMs). Autoregres…

cs.CV2025

ARSS: Taming Decoder-only Autoregressive Visual Generation for View Synthesis From Single View

Wenbin Teng, Gonglin Chen, Haiwei Chen +1

Diffusion models have achieved impressive results in world modeling tasks, including novel view generation from sparse inputs. However, most existing diffusion-based NVS methods ge…

cs.CV2025

FVGen: Accelerating Novel-View Synthesis with Adversarial Video Diffusion Distillation

Wenbin Teng, Gonglin Chen, Haiwei Chen +1

Recent progress in 3D reconstruction has enabled realistic 3D models from dense image captures, yet challenges persist with sparse views, often leading to artifacts in unseen areas…