collaborators

5 papers

cs.DC2025

Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators

Aofeng Shen, Chi Zhang, Yakup Budanaz +4

Tile-based many-Processing Element (PE) accelerators can achieve competitive performance on General Matrix Multiplication (GEMM), but they are extremely hard to program, as their o…

cs.AR2025

FlexLink: Boosting your NVLink Bandwidth by 27% without accuracy concern

Ao Shen, Rui Zhang, Junping Zhao

As large language models (LLMs) continue to scale, multi-node deployment has become a necessity. Consequently, communication has become a critical performance bottleneck. Current i…

cs.CV2025

RadGS-Reg: Registering Spine CT with Biplanar X-rays via Joint 3D Radiative Gaussians Reconstruction and 3D/3D Registration

Ao Shen, Xueming Fu, Junfeng Jiang +6

Computed Tomography (CT)/X-ray registration in image-guided navigation remains challenging because of its stringent requirements for high accuracy and real-time performance. Tradit…

cs.OS2025

Ariadne: A Hotness-Aware and Size-Adaptive Compressed Swap Technique for Fast Application Relaunch and Reduced CPU Usage on Mobile Devices

Yu Liang, Aofeng Shen, Chun Jason Xue +9

Growing application memory demands and concurrent usage are making mobile device memory scarce. When memory pressure is high, current mobile systems use a RAM-based compressed swap…

cs.LG2024

FastSwitch: Optimizing Context Switching Efficiency in Fairness-aware Large Language Model Serving

Ao Shen, Zhiyao Li, Mingyu Gao

Serving numerous users and requests concurrently requires good fairness in Large Language Models (LLMs) serving system. This ensures that, at the same cost, the system can meet the…