2 papers
cs.RO2026
The Gate, Not the Cache: Gate Provenance Bounds the Closed-Loop Reliability of Training-Free VLA Token Skipping
Qi Luo, Shuaijun Liu, Hao Zhao +5
Token skipping is a widely used training-free way to accelerate vision--language--action (VLA) models by bypassing computation for most visual tokens at each control step according…
cs.AI2026
Request-Level Energy Attribution for Batched LLM Serving
Qi Luo, Kunlin Li, Ziwen Wang +2
Batched LLM serving improves throughput but complicates energy accounting. GPU power telemetry is aggregate, whereas sustainability reporting, chargeback, and workload analysis oft…