activity
20242026
collaborators

13 papers

cs.CV2026

Semantic Image Synthesis via Diffusion Models

Wengang Zhou, Weilun Wang, Jianmin Bao +4

Denoising Diffusion Probabilistic Models (DDPMs) have achieved remarkable success in various image generation tasks compared with Generative Adversarial Nets (GANs). Recent work on…

cs.CV2025

Video-based Sign Language Recognition without Temporal Segmentation

Jie Huang, Wengang Zhou, Qilin Zhang +2

Millions of hearing impaired people around the world routinely use some variants of sign languages to communicate, thus the automatic translation of a sign language is meaningful a…

cs.CV2025

SmartEraser: Remove Anything from Images using Masked-Region Guidance

Longtao Jiang, Zhendong Wang, Jianmin Bao +5

Object removal has so far been dominated by the mask-and-inpaint paradigm, where the masked region is excluded from the input, leaving models relying on unmasked areas to inpaint t…

cs.CV2025

VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

Weiye Xu, Jiahao Wang, Weiyun Wang +10

Visual reasoning is a core component of human intelligence and a critical capability for advanced multimodal models. Yet current reasoning evaluations of multimodal large language…

cs.CV2025

Cross-Modal Consistency Learning for Sign Language Recognition

Kepeng Wu, Zecheng Li, Hezhen Hu +2

Pre-training has been proven to be effective in boosting the performance of Isolated Sign Language Recognition (ISLR). Existing pre-training methods solely focus on the compact pos…

cs.CV2025

Uni-Sign: Toward Unified Sign Language Understanding at Scale

Zecheng Li, Wengang Zhou, Weichao Zhao +3

Sign language pre-training has gained increasing attention for its ability to enhance performance across various sign language understanding (SLU) tasks. However, existing methods…