5 papers
An Executable Benchmarking Suite for Tool-Using Agents
Zhiqing Zhong, Zhijing Ye, Jiamin Wang +1
Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-task environments, but benchmark reports often conflate workloads, action-generating dri…
KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving
Zhiqing Zhong, Zhijing Ye, Jian Zhang +3
Static-graph LLM decoders provide predictable launches, fixed tensor shapes, and low submission overhead, but online decoding exposes highly irregular KV-cache behavior: request le…
NCCLZ: Compression-Enabled GPU Collectives with Decoupled Quantization and Entropy Coding
Jiamin Wang, Zhijing Ye, Xiaodong Yu
Collective communication is a major bottleneck for multi-node GPU workloads in scientific computing and distributed deep learning, especially when inter-node bandwidth is limited.…
DABench-LLM: Standardized and In-Depth Benchmarking of Post-Moore Dataflow AI Accelerators for LLMs
Ziyu Hu, Zhiqing Zhong, Weijian Zheng +6
The exponential growth of large language models has outpaced the capabilities of traditional CPU and GPU architectures due to the slowdown of Moore's Law. Dataflow AI accelerators…
An Efficient Gradient-Aware Error-Bounded Lossy Compressor for Federated Learning
Zhijing Ye, Sheng Di, Jiamin Wang +3
Federated learning (FL) enables collaborative model training without exposing clients' private data, but its deployment is often constrained by the communication cost of transmitti…