activity
20242026
collaborators

7 papers

cs.DC2026

Multi-Layer Scheduling for MoE-Based LLM Reasoning

Yifan Sun, Gholamreza Haffari, Minxian Xu +2

Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks, but serving them efficiently at scale remains a critical challenge due to their substant…

cs.CV2026

CoLoGen: Progressive Learning of Concept-Localization Duality for Unified Image Generation

YuXin Song, Yu Lu, Haoyuan Sun +6

Unified conditional image generation remains difficult because different tasks depend on fundamentally different internal representations. Some require conceptual understanding for…

cs.CV2025

Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization

Yang Shen, Xiu-Shen Wei, Yifan Sun +6

Computer Vision (CV) has yet to fully achieve the zero-shot task generalization observed in Natural Language Processing (NLP), despite following many of the milestones established…

cs.CV2024

Dense Connector for MLLMs

Huanjin Yao, Wenhao Wu, Taojiannan Yang +7

Do we fully leverage the potential of visual encoder in Multimodal Large Language Models (MLLMs)? The recent outstanding performance of MLLMs in multimodal understanding has garner…

cs.CV2024

MonoFormer: One Transformer for Both Diffusion and Autoregression

Chuyang Zhao, Yuxing Song, Wenhao Wang +5

Most existing multimodality methods use separate backbones for autoregression-based discrete text generation and diffusion-based continuous visual generation, or the same backbone…

cs.LG2024

Assessing Model Generalization in Vicinity

Yuchi Liu, Yifan Sun, Jingdong Wang +1

This paper evaluates the generalization ability of classification models on out-of-distribution test sets without depending on ground truth labels. Common approaches often calculat…