6 citations · 9 across the 7 of their papers we have counts for
7 papers
LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
Lihao Sun, Hang Dong, Bo Qiao +3
This work characterizes large language models' chain-of-thought generation as a structured trajectory through representation space. We show that mathematical reasoning traverses fu…
An Advanced Reinforcement Learning Framework for Online Scheduling of Deferrable Workloads in Cloud Computing
Hang Dong, Liwen Zhu, Zhao Shan +10
Efficient resource utilization and perfect user experience usually conflict with each other in cloud computing platforms. Great efforts have been invested in increasing resource ut…
Verco: Learning Coordinated Verbal Communication for Multi-agent Reinforcement Learning
Dapeng Li, Hang Dong, Lu Wang +8
In recent years, multi-agent reinforcement learning algorithms have made significant advancements in diverse gaming environments, leading to increased interest in the broader appli…
COIN: Chance-Constrained Imitation Learning for Uncertainty-aware Adaptive Resource Oversubscription Policy
Lu Wang, Mayukh Das, Fangkai Yang +9
We address the challenge of learning safe and robust decision policies in presence of uncertainty in context of the real scientific problem of adaptive resource oversubscription to…
Risk-aware Adaptive Virtual CPU Oversubscription in Microsoft Cloud via Prototypical Human-in-the-loop Imitation Learning
Lu Wang, Mayukh Das, Fangkai Yang +11
Oversubscription is a prevalent practice in cloud services where the system offers more virtual resources, such as virtual cores in virtual machines, to users or applications than…
Counter-Empirical Attacking based on Adversarial Reinforcement Learning for Time-Relevant Scoring System
Xiangguo Sun, Hong Cheng, Hang Dong +3
Scoring systems are commonly seen for platforms in the era of big data. From credit scoring systems in financial services to membership scores in E-commerce shopping platforms, pla…