activity
20242026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

SepPrune:A Separator-based Pruning Framework for Efficient Multimodal Large Language Models

Yuchen Wang, Qihui Zhu, Yang Liu +2

Recent multimodal large language models (MLLMs), such as Qwen2.5-VL and InternVL3, generate large numbers of vision tokens for high-resolution inputs, leading to substantial comput…

cs.CV2026

RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models

Qihui Zhu, Yuchen Wang, Zijian Wen +7

On-Policy Self-Distillation (OPSD) uses privileged information available only to the teacher to provide dense token-level supervision on trajectories generated by the student. Howe…

cs.CV2025

Create Anything Anywhere: Layout-Controllable Personalized Diffusion Model for Multiple Subjects

Wei Li, Hebei Li, Yansong Peng +3

Diffusion models have significantly advanced text-to-image generation, laying the foundation for the development of personalized generative frameworks. However, existing methods la…

cs.CV20241 cited

EE-MLLM: A Data-Efficient and Compute-Efficient Multimodal Large Language Model

Feipeng Ma, Yizhou Zhou, Zheyu Zhang +7

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated satisfactory performance across various vision-language tasks. Current approaches for vision and l…

cs.CV2024

Visual Perception by Large Language Model's Weights

Feipeng Ma, Hongwei Xue, Guangting Wang +7

Existing Multimodal Large Language Models (MLLMs) follow the paradigm that perceives visual information by aligning visual features with the input space of Large Language Models (L…

cs.CV2024

Multi-Modal Generative Embedding Model

Feipeng Ma, Hongwei Xue, Guangting Wang +7

Most multi-modal tasks can be formulated into problems of either generation or embedding. Existing models usually tackle these two types of problems by decoupling language modules…