4 papers
Index SLM Technical Report
Lusheng Zhang, Shien He, Tianxing Yan +8
We present Index-1.9B, a series of open small language models developed at Bilibili. The series comprises four models: Index-1.9B-Base, a foundation model with 1.9 billion non-embe…
When Does Deep RL Beat Calibrated Baselines? A Benchmark Study on Adaptive Resource Control
Guilin Zhang, Chuanyi Sun, Kai Zhao +3
A properly calibrated rule-based autoscaler can beat every one of six mainstream deep reinforcement learning (DRL) algorithms on cost across every workload we test - so when, if ev…
SABER: Switchable and Balanced Training for Efficient LLM Reasoning
Kai Zhao, Yanjun Zhao, Jiaming Song +4
Large language models (LLMs) empowered by chain-of-thought reasoning have achieved impressive accuracy on complex tasks but suffer from excessive inference costs and latency when a…
OPTIMA: Optimized Policy for Intelligent Multi-Agent Systems Enables Coordination-Aware Autonomous Vehicles
Rui Du, Kai Zhao, Jinlong Hou +2
Coordination among connected and autonomous vehicles (CAVs) is advancing due to developments in control and communication technologies. However, much of the current work is based o…