4 papers
SET: Stream-Event-Triggered Scheduling for Efficient CUDA Graph Pipelines
Zhengxiong Li, Tsung-Wei Huang, Umit Ogras
Achieving peak GPU performance remains a significant challenge as the system throughput is constrained by host-device synchronization delays and kernel scheduling overheads, even w…
LEXI: Lossless Exponent Coding for Efficient Inter-Chiplet Communication in Hybrid LLMs
Miao Sun, Alish Kanani, Kaushik Shroff +1
Data movement overheads increase the inference latency of state-of-the-art large language models (LLMs). These models commonly use the bfloat16 (BF16) format for stable training. F…
DUET: Disaggregated Hybrid Mamba-Transformer LLMs with Prefill and Decode-Specific Packages
Alish Kanani, Sangwan Lee, Han Lyu +3
Large language models operate in distinct compute-bound prefill followed by memory bandwidth-bound decode phases. Hybrid Mamba-Transformer models inherit this asymmetry while addin…
K-PACT: Kernel Planning for Adaptive Context Switching -- A Framework for Clustering, Placement, and Prefetching in Spectrum Sensing
H. Umut Suluhan, Jiahao Lin, Serhan Gener +3
Efficient wideband spectrum sensing requires rapid evaluation and re-evaluation of signal presence and type across multiple subchannels. These tasks involve multiple hypothesis tes…