4 papers
Practical One-Round-Trip BFT Replication
Daniel Qian, Xiyu Hao, Jinkun Geng +4
As Byzantine Fault Tolerant (BFT) protocols are increasingly adopted for user-facing applications such as payments and smart contracts, it is crucial that they provide low latency.…
Parallel Prefix Verification for Speculative Generation
Yuncheng Yao, Yuxuan Xia, Shengjie Wang +1
We introduce PARSE (PArallel pRefix Speculative Engine), a speculative generation framework that accelerates large language model (LLM) inference by parallelizing prefix verificati…
Foundry: Template-Based CUDA Graph Context Materialization for Fast LLM Serving Cold Start
Xueshen Liu, Yongji Wu, Yuncheng Yao +3
Modern LLM service providers increasingly rely on autoscaling and parallelism reconfiguration to respond to rapidly changing workloads, but cold-start latency remains a major bottl…
DecodeX: Exploring and Benchmarking of LDPC Decoding across CPU, GPU, and ASIC Platforms
Zhenzhou Qi, Yuncheng Yao, Yiming Li +4
Emerging virtualized radio access networks (vRANs) demand flexible and efficient baseband processing across heterogeneous compute substrates. In this paper, we present DecodeX, a u…