activity
20242026
collaborators

5 papers

cs.CV2026

Lance: Unified Multimodal Modeling by Multi-Task Synergy

Fengyi Fu, Mengqi Huang, Shaojin Wu +10

We present Lance, a lightweight native unified model supporting multimodal understanding, generation, and editing for both images and videos. Rather than relying on model capacity…

cs.CV2025

UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward

Yufeng Cheng, Wenxu Wu, Shaojin Wu +3

Recent advancements in image customization exhibit a wide range of application prospects due to stronger customization capabilities. However, since we humans are more sensitive to…

cs.CV2025

USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning

Shaojin Wu, Mengqi Huang, Yufeng Cheng +5

Existing literature typically treats style-driven and subject-driven generation as two disjoint tasks: the former prioritizes stylistic similarity, whereas the latter insists on su…

cs.CV2025

Less-to-More Generalization: Unlocking More Controllability by In-Context Generation

Shaojin Wu, Mengqi Huang, Wenxu Wu +3

Although subject-driven generation has been extensively explored in image generation due to its wide applications, it still has challenges in data scalability and subject expansibi…

cs.CV2024

VMix: Improving Text-to-Image Diffusion Model with Cross-Attention Mixing Control

Shaojin Wu, Fei Ding, Mengqi Huang +2

While diffusion models show extraordinary talents in text-to-image generation, they may still fail to generate highly aesthetic images. More specifically, there is still a gap betw…