7 papers
SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models
Hongxiang Li, Hongxu Chen, Chenyang Zhu +5
Multimodal Large Language Models (MLLMs) have achieved remarkable success in visual understanding but remain constrained in visual generation due to the fundamental feature discrep…
LISA: Likelihood Score Alignment for Visual-condition Controllable Generation
Yanghao Wang, Hongxu Chen, Jiazhen Liu +4
The prevalent dual-branch paradigm, i.e., training a side network to encode visual conditions and fusing its intermediate-layer features to a frozen pretrained main network, has sh…
Bi-Anchor Interpolation Solver for Accelerating Generative Modeling
Hongxu Chen, Hongxiang Li, Zhen Wang +1
Flow Matching (FM) models have emerged as a leading paradigm for high-fidelity synthesis. However, their reliance on iterative Ordinary Differential Equation (ODE) solving creates…
HMPO: Hybrid Median-length Policy Optimization for Chain-of-Thought Compression
Minghui Zheng, Hongxu Chen, Huimin Ren +8
Large language models achieve remarkable performance via extended chain-of-thought (CoT) reasoning, yet this lengthy process incurs substantial inference overhead. Existing CoT com…
Direct Product Flow Matching: Decoupling Radial and Angular Dynamics for Few-Shot Adaptation
Hongxu Chen, Yanghao Wang, Bowei Zhu +6
Recent flow matching (FM) methods improve the few-shot adaptation of vision-language models, by modeling cross-modal alignment as a continuous multi-step flow. In this paper, we ar…
Zero-Residual Concept Erasure via Progressive Alignment in Text-to-Image Model
Hongxu Chen, Zhen Wang, Taoran Mei +4
Concept Erasure, which aims to prevent pretrained text-to-image models from generating content associated with semantic-harmful concepts (i.e., target concepts), is getting increas…