5 papers
Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index
Outongyi Lv, Yanzhao Zheng, Yuanwei Zhang +5
Reinforcement learning (RL) has become a powerful tool for propelling Large Language Models (LLMs) beyond imitation-based training towards more robust reasoning capabilities. Among…
Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs
Tiancheng Gu, Kaicheng Yang, Ziyong Feng +6
The Contrastive Language-Image Pre-training (CLIP) framework has become a widely used approach for multimodal representation learning, particularly in image-text retrieval and clus…
Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space
Hong Zhang, Zhongjie Duan, Xingjun Wang +6
Unified multimodal generative models aim to integrate image understanding and generation abilities, offering significant advantages in harnessing multimodal corpora, particularly i…
SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning
Yuze Zhao, Jintao Huang, Jinghan Hu +10
Recent development in Large Language Models (LLMs) and Multi-modal Large Language Models (MLLMs) have leverage Attention-based Transformer architectures and achieved superior perfo…
EliGen: Entity-Level Controlled Image Generation with Regional Attention
Hong Zhang, Zhongjie Duan, Xingjun Wang +2
Recent advancements in diffusion models have significantly advanced text-to-image generation, yet global text prompts alone remain insufficient for achieving fine-grained control o…