11 papers
The Gate, Not the Cache: Gate Provenance Bounds the Closed-Loop Reliability of Training-Free VLA Token Skipping
Qi Luo, Shuaijun Liu, Hao Zhao +5
Token skipping is a widely used training-free way to accelerate vision--language--action (VLA) models by bypassing computation for most visual tokens at each control step according…
Request-Level Energy Attribution for Batched LLM Serving
Qi Luo, Kunlin Li, Ziwen Wang +2
Batched LLM serving improves throughput but complicates energy accounting. GPU power telemetry is aggregate, whereas sustainability reporting, chargeback, and workload analysis oft…
ActProbe: Action-Space Probe for Early Failure Detection of Generative Robot Policies
Bingjia Huang, Xiangyu Li, Xiang Wang +7
Generative robot policies fail unpredictably at deployment: they hesitate at critical moments, drift off-task, or commit to unrecoverable actions. Existing online failure detectors…
Unfolding an Atomistic World: Atomistic Simulation of Reactor Pressure Vessel Steel Across Year-and-Meter Scales
Haozhi Han, Ruge Zhang, Haoquan Chen +8
Lifetime prediction of reactor pressure vessel (RPV) steel requires bridging atomistic degradation mechanisms with service-scale spatial and temporal regimes, from Angstroms and pi…
Utility-Aware Data Pricing: Token-Level Quality and Empirical Training Gain for LLMs
Minghui Xu, Qi Luo, Kun Li
Traditional data valuation methods based on ``row-count quality coefficient'' paradigms fail to capture the nuanced, nonlinear contributions that data makes to Large Langu…
M100: An Orchestrated Dataflow Architecture Powering General AI Computing
Yan Xie, Changkui Mao, Changsong Wu +34
As deep learning-based AI technologies gain momentum, the demand for general-purpose AI computing architectures continues to grow. While GPGPU-based architectures offer versatility…