6 papers
Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation
Jiazhen Liu, Mingkuan Feng, Long Chen
MLLM-based segmentation faces a core segmentation trilemma: high segmentation performance, preserved dialogue ability, and fast inference. Embedding-prediction methods may disrupt…
LISA: Likelihood Score Alignment for Visual-condition Controllable Generation
Yanghao Wang, Hongxu Chen, Jiazhen Liu +4
The prevalent dual-branch paradigm, i.e., training a side network to encode visual conditions and fusing its intermediate-layer features to a frozen pretrained main network, has sh…
DepCap: Adaptive Block-Wise Parallel Decoding for Efficient Diffusion LM Inference
Xiang Xia, Wuyang Zhang, Jiazheng Liu +2
Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive language generation due to their potential for parallel decoding and global refinement of…
Empowering Small VLMs to Think with Dynamic Memorization and Exploration
Jiazhen Liu, Yuchuan Deng, Long Chen
Small-scale Vision-Language Models (SVLMs) are exceptionally well-suited for proprietary tasks. Equipping them with thinking capabilities is a critical step to enhance their perfor…
Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction
Jiazhen Liu, Mingkuan Feng, Long Chen
Integrating segmentation into Multimodal Large Language Models (MLLMs) presents a core trilemma: simultaneously preserving dialogue ability, achieving high segmentation performance…
Segmentation as A Plug-and-Play Capability for Frozen Multimodal LLMs
Jiazhen Liu, Long Chen
Integrating diverse visual capabilities into a unified model is a significant trend in Multimodal Large Language Models (MLLMs). Among these, the inclusion of segmentation poses a…