activity
20222026
most citedDisentangled Generation with Information Bottleneck for Few-Shot Learning

1 citations · 1 across the 6 of their papers we have counts for

collaborators

18 papers

cs.CV2026

CameraEditor: Camera-Controlled Image Editing via Video-Prior Sequential Modeling

Xin Shen, Chengyou Jia, Keshuo Xing +6

Beyond semantic content, camera parameters play a pivotal role in dictating the geometric perspective and appearance of any given image. While recent image editing models excel at…

cs.CV2026

OmniTryOn: Video Try-On Anything at Once!

Changliang Xia, Chengyou Jia, Minnan Luo +3

Although video virtual try-on (VVT) has achieved significant progress, existing methods still exhibit two fundamental limitations: first, they are restricted to single-garment tran…

cs.CV2026

-Predictor: Noise-Free Deterministic Diffusion for Dense Prediction

Changliang Xia, Chengyou Jia, Minnan Luo +3

Although diffusion models with strong visual priors have emerged as powerful dense prediction backbones, they overlook a core limitation: the stochastic noise at the core of diffus…

cs.CV2025

PaCo-RL: Advancing Reinforcement Learning for Consistent Image Generation with Pairwise Reward Modeling

Bowen Ping, Chengyou Jia, Minnan Luo +4

Consistent image generation requires faithfully preserving identities, styles, and logical coherence across multiple images, which is essential for applications such as storytellin…

cs.CV2025

CoFFT: Chain of Foresight-Focus Thought for Visual Language Models

Xinyu Zhang, Yuxuan Dong, Lingling Zhang +5

Despite significant advances in Vision Language Models (VLMs), they remain constrained by the complexity and redundancy of visual input. When images contain large amounts of irrele…

cs.CV2025

Multi-Modal Dataset Distillation in the Wild

Zhuohang Dang, Minnan Luo, Chengyou Jia +3

Recent multi-modal models have shown remarkable versatility in real-world applications. However, their rapid development encounters two critical data challenges. First, the trainin…