7 papers
SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning
Zhen-Hao Xie, Jun-Tao Tang, Yu-Cheng Shi +3
Multimodal Large Language Models (MLLMs) achieve strong performance through instruction tuning, but real-world deployment requires them to continually expand their capabilities, ma…
Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling
Xiaolong Fu, Lichen Ma, Zipeng Guo +9
The integration of Reinforcement Learning (RL) into flow matching models for text-to-image (T2I) generation has driven substantial advances in generation quality. However, these ga…
UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing
Lichen Ma, Xiaolong Fu, Gaojing Zhou +6
With the rapid advancement of image generation, visual text editing using natural language instructions has received increasing attention. The main challenge of this task is to ful…
HYDRA: Unifying Multi-modal Generation and Understanding via Representation-Harmonized Tokenization
Xuerui Qiu, Yutao Cui, Guozhen Zhang +9
Unified Multimodal Models struggle to bridge the fundamental gap between the abstract representations needed for visual understanding and the detailed primitives required for gener…
Think Bright, Diffuse Nice: Enhancing T2I-ICL via Inductive-Bias Hint Instruction and Query Contrastive Decoding
Zhiyong Ma, Zhenpeng Li, Yuanjie Shi +3
Text-to-Image In-Context Learning (T2I-ICL) enables customized image synthesis via interleaved text-image examples but faces two mutually reinforcing bottlenecks, compliance failur…
DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion Representations
Yuxiang Shi, Zhe Li, Yanwen Wang +3
Portrait animation from a single source image and a driving video is a long-standing problem. Recent approaches tend to adopt diffusion-based image/video generation models for real…