activity
20242026
most citedMDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture Generation

6 citations · 21 across the 41 of their papers we have counts for

collaborators

59 papers

cs.CV2026

Evidence-RL: Towards Evidence-intensive Visual Reasoning

Haojie Huang, Xinlei Yu, Chengming Xu +6

Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware pos…

cs.CV2026

In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusion

Lingxiao Yang, Liu Liu, Moran Li +4

Current few-step autoregressive video diffusion models depend on previous fully denoised clean frames as context for all denoising steps of the current frame. However, these clean…

cs.CV2026

MambaADv2: Evolving Duality-enhanced State Space Model for Unsupervised Anomaly Detection

Xiaobin Hu, Haoyang He, Bo Yin +5

While recent advancements in anomaly detection have demonstrated the efficacy of CNN- and Transformer-based approaches, these architectures face inherent limitations: CNNs struggle…

cs.AI2026

Dual Latent Memory for Visual Multi-agent System

Xinlei Yu, Chengming Xu, Zhangquan Chen +8

While Visual Multi-Agent Systems (VMAS) promise to enhance comprehensive abilities through inter-agent collaboration, empirical evidence reveals a counter-intuitive "scaling wall":…

cs.CV2026

FFP-300K: Scaling First-Frame Propagation for Generalizable Video Editing

Xijie Huang, Chengming Xu, Donghao Luo +6

First-Frame Propagation (FFP) offers a promising paradigm for controllable video editing, but existing methods are hampered by a reliance on cumbersome run-time guidance. We identi…

cs.CV2026

Towards Generalized Multi-Image Editing for Unified Multimodal Models

Pengcheng Xu, Peng Tang, Donghao Luo +7

Unified Multimodal Models (UMMs) integrate multimodal understanding and generation, yet they are limited to maintaining visual consistency and disambiguating visual cues when refer…