4 papers
VoltanaLLM: Energy-Efficient and SLO-Aware Disaggregated LLM Serving via Adaptive Frequency Control and State-Space Routing
Jiahuan Yu, Aryan Taneja, Junfeng Lin +1
The energy cost of Large Language Model (LLM) inference is rapidly becoming a barrier to sustainable and scalable deployment. Although modern serving architectures expose distinct…
User Prompting Strategies and Prompt Enhancement Methods for Open-Set Object Detection in XR Environments
Junfeng Lin, Yanming Xiu, Maria Gorlatova
Open-set object detection (OSOD) localizes objects while identifying and rejecting unknown classes at inference. While recent OSOD models perform well on benchmarks, their behavior…
CompAir: Synergizing Complementary PIMs and In-Transit NoC Computation for Efficient LLM Acceleration
Hongyi Li, Songchen Ma, Huanyu Qu +5
The rapid advancement of Large Language Models (LLMs) has revolutionized various aspects of human life, yet their immense computational and energy demands pose significant challeng…
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
Huanyu Qu, Weihao Zhang, Junfeng Lin +4
To efficiently support large-scale NNs, multi-level hardware, leveraging advanced integration and interconnection technologies, has emerged as a promising solution to counter the s…