4 papers
Efficient Multi-round LLM Inference over Disaggregated Serving
Wenhao He, Youhe Jiang, Penghao Zhao +4
With the rapid evolution of Large Language Models (LLMs), multi-round workflows, such as autonomous agents and iterative retrieval, have become increasingly prevalent. However, thi…
MixServe: An Automatic Distributed Serving System for MoE Models with Hybrid Parallelism Based on Fused Communication Algorithm
Bowen Zhou, Jinrui Jia, Wenhao He +2
The Mixture of Experts (MoE) models are emerging as the latest paradigm for Large Language Models (LLMs). However, due to memory constraints, MoE models with billions or even trill…
Exponential quantum advantages for practical non-Hermitian eigenproblems
Xiao-Ming Zhang, Yukun Zhang, Wenhao He +1
Non-Hermitian physics has emerged as a rich field of study, with applications ranging from -symmetry breaking and skin effects to non-Hermitian topological phase transitions. Y…
SAM-IF: Leveraging SAM for Incremental Few-Shot Instance Segmentation
Xudong Zhou, Wenhao He
We propose SAM-IF, a novel method for incremental few-shot instance segmentation leveraging the Segment Anything Model (SAM). SAM-IF addresses the challenges of class-agnostic inst…