3 papers
cs.DC2026
PipeSwift: Revisiting Pipeline Parallelism for Large-Scale Completion-Oriented Agentic Serving
Shiju Wang, Fei Ren, Fangcheng Fu +4
LLM agents execute long-horizon workflows where each model response determines the progress of subsequent tool interactions and environment transitions. Unlike chatbot serving, whe…
cs.AR2025
SoMa: Identifying, Exploring, and Understanding the DRAM Communication Scheduling Space for DNN Accelerators
Jingwei Cai, Xuan Wang, Mingyu Gao +5
Modern Deep Neural Network (DNN) accelerators are equipped with increasingly larger on-chip buffers to provide more opportunities to alleviate the increasingly severe DRAM bandwidt…
cs.AR2023
Gemini: Mapping and Architecture Co-exploration for Large-scale DNN Chiplet Accelerators
Jingwei Cai, Zuotong Wu, Sen Peng +5
Chiplet technology enables the integration of an increasing number of transistors on a single accelerator with higher yield in the post-Moore era, addressing the immense computatio…