2 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.CL2025
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving
Ying Wang, Zhen Jin, Jiexiong Xu +3
As augmented large language models (LLMs) with external tools become increasingly popular in web applications, improving augmented LLM inference serving efficiency and optimizing s…
cs.OS2025
GoCkpt: Gradient-Assisted Multi-Step overlapped Checkpointing for Efficient LLM Training
Keyao Zhang, Yiquan Chen, Zhuo Hu +3
The accuracy of large language models (LLMs) improves with increasing model size, but increasing model complexity also poses significant challenges to training stability. Periodic…
cs.AR2023★ 2 cited
High-performance and Scalable Software-based NVMe Virtualization Mechanism with I/O Queues Passthrough
Yiquan Chen, Zhen Jin, Yijing Wang +10
NVMe(Non-Volatile Memory Express) is an industry standard for solid-state drives (SSDs) that has been widely adopted in data centers. NVMe virtualization is crucial in cloud comput…