3 papers
cs.SE2026
EnerInfer: Energy-Aware On-Device LLM Inference
Bohua Zou, Nian Liu, Binqi Sun +6
On-device LLM inference is increasingly attractive for privacy-preserving, reliable, and cost-effective deployment, yet its energy and thermal costs remain a critical bottleneck. E…
cs.SE2026
ProfInfer: An eBPF-based Fine-Grained LLM Inference Profiler
Bohua Zou, Debayan Roy, Dhimankumar Yogesh Airao +4
As large language models (LLMs) move from research to production, understanding how inference engines behave in real time has become both essential and elusive. Unlike general-purp…
eess.SP2025
Profiling Multi-Level Operator Costs for Bottleneck Diagnosis in High-Speed Data Planes
Zhiyuan Ren, Yutao Liu, Wenchi Cheng +1
This paper proposes a saturation throughput delta-based methodology to precisely measure operator costs in high-speed data planes without intrusive instrumentation. The approach ca…