activity
20242026
collaborators

23 papers

cs.CV2026

MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving

Zhijing Cheng, Xuancheng Zhang, Donglin Di +4

End-to-end autonomous driving systems commonly follow a cascaded two-stage pipeline where a perception stage compresses multi-modal sensor inputs into a compact context and a downs…

cs.CV2026

Straight-Path Flow Matching for Incomplete Multi-View Clustering

Yiteng Yuan, Junyan Wang, Zheyuan Liu +4

Incomplete Multi-View Clustering addresses the problem of clustering multi-modal data when certain views are missing. Recent end-to-end generative approaches leverage diffusion mod…

cs.CV2026

ArcAD: Anomaly-Rectified Calibration for Cold-Start Supervised Anomaly Detection

Ningning Han, Lei Fan, Jia Guo +5

The deployment of Industrial Anomaly Detection (IAD) in real-world manufacturing frequently encounters a challenging cold-start bottleneck, in which limited normal samples fail to…

cs.CV2026

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning

He Feng, Yongjia Ma, Donglin Di +2

Portrait animation methods have achieved substantial visual quality and lip synchronization, but fine-grained manipulation of the eye region still faces a trade-off between input g…

cs.CV2026

Chain of World: World Model Thinking in Latent Motion

Fuxiang Yang, Donglin Di, Lulu Tang +6

Vision-Language-Action (VLA) models are a promising path toward embodied intelligence, yet they often overlook the predictive and temporal-causal structure underlying visual dynami…

cs.CV2026

Visual Prompt-Agnostic Evolution

Junze Wang, Lei Fan, Dezheng Zhang +5

Visual Prompt Tuning (VPT) adapts a frozen Vision Transformer (ViT) to downstream tasks by inserting a small number of learnable prompt tokens into the token sequence at each layer…