From the 1 of 7 linked papers with an AI index.
7 papers
SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer
Xuchuan Luo, Jiacheng Shen, Xin Wang +1
The paper introduces SmartGen, a system that reduces network overhead in disaggregated large language model inference by selectively transferring only essential key‑value cache ent…
Bringing Managed Language Support to WebAssembly with External Library Linking
Shuyao Jiang, Ruiying Zeng, Yangfan Zhou +1
WebAssembly (Wasm) has emerged as a powerful bytecode format for running applications with near-native performance in portable and secure environments. However, while Wasm currentl…
Debugging Performance Issues in WebAssembly Runtimes via Mutation-based Inference
Ruiying Zeng, Shuyao Jiang, Wenxuan Zhao +1
Performance debugging in WebAssembly (Wasm) runtimes is essential for ensuring the robustness of Wasm, especially since performance issues have frequently occurred in Wasm runtimes…
EdgeFlow: Fast Cold Starts for LLMs on Mobile Devices
Yongsheng Yan, Jiacheng Shen, Xuchuan Luo +1
Deploying large language models (LLMs) on mobile devices is an emerging trend to enable data privacy and offline accessibility of LLM applications. Modern mobile neural processing…
Can User Feedback Help Issue Detection? An Empirical Study on a One-billion-user Online Service System
Shuyao Jiang, Jiazhen Gu, Wujie Zheng +2
Background: It has long been suggested that user feedback, typically written in natural language by end-users, can help issue detection. However, for large-scale online service sys…
Distinguishability-guided Test Program Generation for WebAssembly Runtime Performance Testing
Shuyao Jiang, Ruiying Zeng, Yangfan Zhou +1
WebAssembly (Wasm) is a binary instruction format designed as a portable compilation target, which has been widely used on both the web and server sides in recent years. As high pe…