2 papers
cs.DC2026
Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving
Jianru Ding, Ryien Hosseini, Pouya Mahdi Gholami +2
LLM-based agents resolve a user task through many turns of dependent inference and tool calls, producing a workload whose total cost is unknown when the task arrives. Existing mult…
cs.ET2024
Drop-Connect as a Fault-Tolerance Approach for RRAM-based Deep Neural Network Accelerators
Mingyuan Xiang, Xuhan Xie, Pedro Savarese +3
Resistive random-access memory (RRAM) is widely recognized as a promising emerging hardware platform for deep neural networks (DNNs). Yet, due to manufacturing limitations, current…