5 papers
Rethinking the Evaluation of Efficiency Methods for Multi-Agent Systems
Jiamu Zhang, Lingxi Zhang, Pengjun Lu +6
Efficiency is increasingly important for Large Language Model (LLM)-based multi-agent systems (MAS), as larger models and more agents introduce substantial execution costs. Recent…
SPECTRA: Pushing the KV Cache Beyond the 2-Bit Cliff via Spectral Transform Coding
Jiamu Zhang, Liang Wu, Kelly Wan +2
Large language models (LLMs) increasingly read long inputs in the agentic era, from whole documents and codebases to conversations across many turns. Their inference memory is then…
WiSP: A Working-Set View of Mixture-of-Experts Serving on Extremely Low-Resource Hardware
Jiamu Zhang, Liang Wu, Mayank Darbari +1
Modern local and agentic workloads often need large-model capacity at low concurrency, but run on GPUs that cannot keep a frontier-scale model resident. Mixture-of-Experts (MoE) mo…
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Yang Sui, Yu-Neng Chuang, Guanchu Wang +9
Large Language Models (LLMs) have demonstrated remarkable capabilities in complex tasks. Recent advancements in Large Reasoning Models (LRMs), such as OpenAI o1 and DeepSeek-R1, ha…
The Science of Evaluating Foundation Models
Jiayi Yuan, Jiamu Zhang, Andrew Wen +1
The emergent phenomena of large foundation models have revolutionized natural language processing. However, evaluating these models presents significant challenges due to their siz…