2 papers
cs.DC2025
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
Chenxiang Ma, Zhisheng Ye, Hanyu Zhao +9
Offloading large language models (LLMs) state to host memory during inference promises to reduce operational costs by supporting larger models, longer inputs, and larger batch size…
cs.SE2024
LEADS: Lightweight Embedded Assisted Driving System
Tianhao Fu, Querobin Mascarenhas, Andrew Forti
With the rapid development of electric vehicles, formula races that face high school and university students have become more popular than ever as the threshold for design and manu…