collaborators

5 papers

cs.CV2026

RubricRL: Simple Generalizable Rewards for Text-to-Image Generation

Xuelu Feng, Yunsheng Li, Ziyu Wan +4

Reinforcement learning (RL) has recently emerged as a promising approach for aligning text-to-image generative models with human preferences. A key challenge, however, lies in desi…

cs.CV2026

Leveraging Text-to-Image Diffusion Models for Unsupervised Visual Object Tracking

Zhengbo Zhang, Zhigang Tu, Junsong Yuan +2

Unsupervised visual object tracking is a challenging task that requires following arbitrary targets in videos without training on ground-truth annotations. Despite considerable pro…

cs.CV2025

GeoRemover: Removing Objects and Their Causal Visual Artifacts

Zixin Zhu, Haoxiang Li, Xuelu Feng +3

Towards intelligent image editing, object removal should eliminate both the target object and its causal visual artifacts, such as shadows and reflections. However, existing image…

cs.CV2025

CompSlider: Compositional Slider for Disentangled Multiple-Attribute Image Generation

Zixin Zhu, Kevin Duarte, Mamshad Nayeem Rizve +3

In text-to-image (T2I) generation, achieving fine-grained control over attributes - such as age or smile - remains challenging, even with detailed text prompts. Slider-based method…

cs.CV2025

Benchmarking Large and Small MLLMs

Xuelu Feng, Yunsheng Li, Dongdong Chen +4

Large multimodal language models (MLLMs) such as GPT-4V and GPT-4o have achieved remarkable advancements in understanding and generating multimodal content, showcasing superior qua…