4 papers
Libra: Efficient Resource Management for Agentic RL Post-Training
Kaiwen Chen, Xin Tan, Jingzong Li +1
Reinforcement learning (RL) has emerged as a standard post-training paradigm for shaping large language models (LLMs) into capable agents. In agentic RL, the rollout stage generate…
ReasonCache: Accelerating Large Reasoning Model Serving through KV Cache Sharing
Kaiwen Chen, Xin Tan, Minchen Yu +2
Large Reasoning Models (LRMs) are becoming integral to many AI inference systems, enhancing their capabilities with advanced reasoning. However, deploying these models in productio…
Easz: An Agile Transformer-based Image Compression Framework for Resource-constrained IoTs
Yu Mao, Jingzong Li, Jun Wang +4
Neural image compression, necessary in various machine-to-machine communication scenarios, suffers from its heavy encode-decode structures and inflexibility in switching between di…
Echo: Simulating Distributed Training At Scale
Yicheng Feng, Yuetao Chen, Kaiwen Chen +7
Simulation offers unique values for both enumeration and extrapolation purposes, and is becoming increasingly important for managing the massive machine learning (ML) clusters and…