activity
20242026
collaborators

11 papers

cs.CV2026

Cross-Layer Attentive Feature Upsampling for Low-latency Semantic Segmentation

Tianheng Cheng, Xinggang Wang, Junchao Liao +1

Semantic segmentation is a fundamental problem in computer vision and it requires high-resolution feature maps for dense prediction. Current coordinate-guided low-resolution featur…

cs.CV2025

4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer

Xianfeng Wu, Yajing Bai, Minghan Li +5

Constructing 4D language fields is crucial for embodied AI, augmented/virtual reality, and 4D scene understanding, as they provide enriched semantic representations of dynamic envi…

cs.CV2025

Visual Generation Tuning

Jiahao Guo, Sinan Du, Jingfeng Yao +7

Large Vision Language Models (VLMs) effectively bridge the modality gap through extensive pretraining, acquiring sophisticated visual representations aligned with language. However…

cs.CV2025

Gait Recognition via Collaborating Discriminative and Generative Diffusion Models

Haijun Xiong, Bin Feng, Bang Wang +2

Gait recognition offers a non-intrusive biometric solution by identifying individuals through their walking patterns. Although discriminative models have achieved notable success i…

cs.CV2025

Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile Devices

Ya Zou, Jingfeng Yao, Siyuan Yu +3

There is a growing demand for deploying large generative AI models on mobile devices. For recent popular video generative models, however, the Variational AutoEncoder (VAE) represe…

cs.CV2025

PixelHacker: Image Inpainting with Structural and Semantic Consistency

Ziyang Xu, Kangsheng Duan, Xiaolei Shen +5

Image inpainting is a fundamental research area between image editing and image generation. Recent state-of-the-art (SOTA) methods have explored novel attention mechanisms, lightwe…