3 papers
cs.LG2026
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning
Zhaopeng Qiu, Shuang Yu, Jingqi Zhang +4
Reinforcement learning (RL) for large language models (LLMs) is increasingly bottlenecked by rollout (generation), where long output sequence lengths make attention and KV-cache me…
cs.LG2026
AIConfigurator: Lightning-Fast Configuration Optimization for Multi-Framework LLM Serving
Tianhao Xu, Yiming Liu, Xianglong Lu +18
Optimizing Large Language Model (LLM) inference in production systems is increasingly difficult due to dynamic workloads, stringent latency/throughput targets, and a rapidly expand…
cs.DC2025
DTVM: Revolutionizing Smart Contract Execution with Determinism and Compatibility
Wei Zhou, Xiong Xu, Changzheng Wei +22
We introduce the DeTerministic Virtual Machine (DTVM) Stack, a next-generation smart contract execution framework designed to address critical performance, determinism, and ecosyst…