3 papers
cs.DC2025
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints
Yichao Yuan, Lin Ma, Nishil Talati
Mixture of Experts (MoE) LLMs, characterized by their sparse activation patterns, offer a promising approach to scaling language models while avoiding proportionally increasing the…
cs.DB2025
Vortex: Overcoming Memory Capacity Limitations in GPU-Accelerated Large-Scale Data Analytics
Yichao Yuan, Advait Iyer, Lin Ma +1
Despite the high computational throughput of GPUs, limited memory capacity and bandwidth-limited CPU-GPU communication via PCIe links remain significant bottlenecks for acceleratin…
cs.DC2021
Maxwell: a hardware and software highly integrated compute-storage system
ing Ma, Wei Zhou, Sihai Yao +4
The compute-storage framework is responsible for data storage and processing, and acts as the digital chassis of all upper-level businesses. The performance of the framework affect…