2 papers
cs.CV2025
LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models
Juntao Liu, Liqiang Niu, Wenchao Chen +2
Existing visual token compression methods for Multimodal Large Language Models (MLLMs) predominantly operate as post-encoder modules, limiting their potential for efficiency gains.…
cs.CV2024
MaskMamba: A Hybrid Mamba-Transformer Model for Masked Image Generation
Wenchao Chen, Liqiang Niu, Ziyao Lu +2
Image generation models have encountered challenges related to scalability and quadratic complexity, primarily due to the reliance on Transformer-based backbones. In this study, we…