5 papers
SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models
Hongxiang Li, Hongxu Chen, Chenyang Zhu +5
Multimodal Large Language Models (MLLMs) have achieved remarkable success in visual understanding but remain constrained in visual generation due to the fundamental feature discrep…
Direct Product Flow Matching: Decoupling Radial and Angular Dynamics for Few-Shot Adaptation
Hongxu Chen, Yanghao Wang, Bowei Zhu +6
Recent flow matching (FM) methods improve the few-shot adaptation of vision-language models, by modeling cross-modal alignment as a continuous multi-step flow. In this paper, we ar…
ASK: Adaptive Self-improving Knowledge Framework for Audio Text Retrieval
Siyuan Fu, Xuchen Guo, Mingjun Liu +7
The dominant paradigm for Audio-Text Retrieval (ATR) relies on dual-encoder architectures optimized via mini-batch contrastive learning. However, restricting optimization to local…
MoKus: Leveraging Cross-Modal Knowledge Transfer for Knowledge-Aware Concept Customization
Chenyang Zhu, Hongxiang Li, Xiu Li +1
Concept customization typically binds rare tokens to a target concept. Unfortunately, these approaches often suffer from unstable performance as the pretraining data seldom contain…
Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine
Xiaoshuang Huang, Lingdong Shen, Jia Liu +4
In recent years, Multimodal Large Language Models (MLLM) have achieved notable advancements, demonstrating the feasibility of developing an intelligent biomedical assistant. Howeve…