collaborators

7 papers

cs.CV2026

Object-Centric Vision Token Pruning for Vision Language Models

Guangyuan Li, Rongzhen Zhao, Jinhong Deng +2

In Vision Language Models (VLMs), vision tokens are quantity-heavy yet information-dispersed compared with language tokens, thus consume too much unnecessary computation. Pruning r…

cs.CV2026

SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing

Ying Zeng, Miaosen Luo, Guangyuan Li +10

Traditional photographic image editing typically requires users to possess sufficient aesthetic understanding to provide appropriate instructions for adjusting image quality and ca…

cs.CV2026

MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration

Guangyuan Li, Bo Li, Jinwei Chen +3

Recent interactive video world model methods generate scene evolution conditioned on user instructions. Although they achieve impressive results, two key limitations remain. First,…

cs.CV2025

MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on

Guangyuan Li, Siming Zheng, Hao Zhang +6

Video Virtual Try-On (VVT) aims to synthesize garments that appear natural across consecutive video frames, capturing both their dynamics and interactions with human motion. Despit…

cs.CV2025

Incomplete Multi-view Clustering via Diffusion Contrastive Generation

Yuanyang Zhang, Yijie Lin, Weiqing Yan +6

Incomplete multi-view clustering (IMVC) has garnered increasing attention in recent years due to the common issue of missing data in multi-view datasets. The primary approach to ad…

cs.CV2025

DyArtbank: Diverse Artistic Style Transfer via Pre-trained Stable Diffusion and Dynamic Style Prompt Artbank

Zhanjie Zhang, Quanwei Zhang, Guangyuan Li +4

Artistic style transfer aims to transfer the learned style onto an arbitrary content image. However, most existing style transfer methods can only render consistent artistic styliz…