3 papers
cs.CV2025
Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space
Hong Zhang, Zhongjie Duan, Xingjun Wang +6
Unified multimodal generative models aim to integrate image understanding and generation abilities, offering significant advantages in harnessing multimodal corpora, particularly i…
cs.CL2024
Minimum Tuning to Unlock Long Output from LLMs with High Quality Data as the Key
Yingda Chen, Xingjun Wang, Jintao Huang +3
As large language models rapidly evolve to support longer context, there is a notable disparity in their capability to generate output at greater lengths. Recent study suggests tha…
cs.CL2024
SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning
Yuze Zhao, Jintao Huang, Jinghan Hu +10
Recent development in Large Language Models (LLMs) and Multi-modal Large Language Models (MLLMs) have leverage Attention-based Transformer architectures and achieved superior perfo…