1 paper
Qi Luo, Kunlin Li, Ziwen Wang +2
Batched LLM serving improves throughput but complicates energy accounting. GPU power telemetry is aggregate, whereas sustainability reporting, chargeback, and workload analysis oft…