collaborators

5 papers

cs.AI2026

Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index

Outongyi Lv, Yanzhao Zheng, Yuanwei Zhang +5

Reinforcement learning (RL) has become a powerful tool for propelling Large Language Models (LLMs) beyond imitation-based training towards more robust reasoning capabilities. Among…

cs.CV2025

Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs

Tiancheng Gu, Kaicheng Yang, Ziyong Feng +6

The Contrastive Language-Image Pre-training (CLIP) framework has become a widely used approach for multimodal representation learning, particularly in image-text retrieval and clus…

cs.CV2025

Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space

Hong Zhang, Zhongjie Duan, Xingjun Wang +6

Unified multimodal generative models aim to integrate image understanding and generation abilities, offering significant advantages in harnessing multimodal corpora, particularly i…

cs.CL2025

SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning

Yuze Zhao, Jintao Huang, Jinghan Hu +10

Recent development in Large Language Models (LLMs) and Multi-modal Large Language Models (MLLMs) have leverage Attention-based Transformer architectures and achieved superior perfo…

cs.CV2025

EliGen: Entity-Level Controlled Image Generation with Regional Attention

Hong Zhang, Zhongjie Duan, Xingjun Wang +2

Recent advancements in diffusion models have significantly advanced text-to-image generation, yet global text prompts alone remain insufficient for achieving fine-grained control o…