7 papers
Multi-Layer Scheduling for MoE-Based LLM Reasoning
Yifan Sun, Gholamreza Haffari, Minxian Xu +2
Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks, but serving them efficiently at scale remains a critical challenge due to their substant…
CoLoGen: Progressive Learning of Concept-Localization Duality for Unified Image Generation
YuXin Song, Yu Lu, Haoyuan Sun +6
Unified conditional image generation remains difficult because different tasks depend on fundamentally different internal representations. Some require conceptual understanding for…
Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization
Yang Shen, Xiu-Shen Wei, Yifan Sun +6
Computer Vision (CV) has yet to fully achieve the zero-shot task generalization observed in Natural Language Processing (NLP), despite following many of the milestones established…
Dense Connector for MLLMs
Huanjin Yao, Wenhao Wu, Taojiannan Yang +7
Do we fully leverage the potential of visual encoder in Multimodal Large Language Models (MLLMs)? The recent outstanding performance of MLLMs in multimodal understanding has garner…
MonoFormer: One Transformer for Both Diffusion and Autoregression
Chuyang Zhao, Yuxing Song, Wenhao Wang +5
Most existing multimodality methods use separate backbones for autoregression-based discrete text generation and diffusion-based continuous visual generation, or the same backbone…
Assessing Model Generalization in Vicinity
Yuchi Liu, Yifan Sun, Jingdong Wang +1
This paper evaluates the generalization ability of classification models on out-of-distribution test sets without depending on ground truth labels. Common approaches often calculat…