activity
20242026
collaborators

5 papers

cs.CV2026

PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards

Shulei Wang, Longhui Wei, Xin He +4

Personalized generation models for a single subject have demonstrated remarkable effectiveness, highlighting their significant potential. However, when extended to multiple subject…

cs.CV2025

EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture

Xin He, Longhui Wei, Jianbo Ouyang +3

We propose EMMA, an efficient and unified architecture for multimodal understanding, generation and editing. Specifically, EMMA primarily consists of 1) An efficient autoencoder wi…

cs.CV2025

Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts

Xin He, Xumeng Han, Longhui Wei +2

Multimodal large language models (MLLMs) require a nuanced interpretation of complex image information, typically leveraging a vision encoder to perceive various visual scenarios.…

cs.CV2025

Boosting Segment Anything Model Towards Open-Vocabulary Learning

Xumeng Han, Longhui Wei, Xuehui Yu +6

The recent Segment Anything Model (SAM) has emerged as a new paradigmatic vision foundation model, showcasing potent zero-shot generalization and flexible prompting. Despite SAM fi…

cs.CV2024

ViMoE: An Empirical Study of Designing Vision Mixture-of-Experts

Xumeng Han, Longhui Wei, Zhiyang Dou +6

Mixture-of-Experts (MoE) models embody the divide-and-conquer concept and are a promising approach for increasing model capacity, demonstrating excellent scalability across multipl…