3 papers
cs.LG2026
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
Feiyu Yao, Zhixiong Niu, Xiaqing Li +3
Long-context inference increasingly operates over CPU-resident KV caches, either because decoding-time KV states exceed GPU memory capacity or because disaggregated prefill-decode…
cs.LG2025
NestQuant: Post-Training Integer-Nesting Quantization for On-Device DNN
Jianhang Xie, Chuntao Ding, Xiaqing Li +3
Deploying quantized deep neural network (DNN) models with resource adaptation capabilities on ubiquitous Internet of Things (IoT) devices to provide high-quality AI services can le…
cs.AR2025
AGON: Automated Design Framework for Customizing Processors from ISA Documents
Chongxiao Li, Di Huang, Pengwei Jin +12
Customized processors are attractive solutions for vast domain-specific applications due to their high energy efficiency. However, designing a processor in traditional flows is tim…