3 papers
cs.CV2025
Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space
Hong Zhang, Zhongjie Duan, Xingjun Wang +6
Unified multimodal generative models aim to integrate image understanding and generation abilities, offering significant advantages in harnessing multimodal corpora, particularly i…
cs.CL2025
SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning
Yuze Zhao, Jintao Huang, Jinghan Hu +10
Recent development in Large Language Models (LLMs) and Multi-modal Large Language Models (MLLMs) have leverage Attention-based Transformer architectures and achieved superior perfo…
cs.CL2024
Minimum Tuning to Unlock Long Output from LLMs with High Quality Data as the Key
Yingda Chen, Xingjun Wang, Jintao Huang +3
As large language models rapidly evolve to support longer context, there is a notable disparity in their capability to generate output at greater lengths. Recent study suggests tha…