papers

Publications (8)

cs.LG2024

Value-Based Deep Multi-Agent Reinforcement Learning with Dynamic Sparse Training

Pihe Hu, Shaolong Li, Zhuoran Li +2

Deep Multi-agent Reinforcement Learning (MARL) relies on neural networks with numerous parameters in multi-agent scenarios, often incurring substantial computational overhead. Cons…

cs.IT2018

Optimal Hybrid Full-Duplex/Half-Duplex Scheme for Buffer Aided Relay Systems

Cheng Li, Bin Xia, Pihe Hu +1

Full-duplex (FD) communication has received great interest in recent years due to the potential of doubling the spectral efficiency. However, how to alleviate the detrimental effec…

cs.LG2023

Nearly Minimax Optimal Reinforcement Learning with Linear Function Approximation

Pihe Hu, Yu Chen, Longbo Huang

We study reinforcement learning with linear function approximation where the transition probability and reward functions are linear with respect to a feature mapping $\boldsymbolϕ…

cs.LG2022

Effective Multi-User Delay-Constrained Scheduling with Deep Recurrent Reinforcement Learning

Pihe Hu, Ling Pan, Yu Chen +2

Multi-user delay constrained scheduling is important in many real-world applications including wireless communication, live streaming, and cloud computing. Yet, it poses a critical…

cs.LG2024

Mixed Sparsity Training: Achieving 4 FLOP Reduction for Transformer Pretraining

Pihe Hu, Shaolong Li, Longbo Huang

Large language models (LLMs) have made significant strides in complex tasks, yet their widespread adoption is impeded by substantial computational demands. With hundreds of billion…

cs.LG2023

RLx2: Training a Sparse Deep Reinforcement Learning Model from Scratch

Yiqin Tan, Pihe Hu, Ling Pan +2

Training deep reinforcement learning (DRL) models usually requires high computation costs. Therefore, compressing DRL models possesses immense potential for training acceleration a…

cs.IT2018

Optimal Multi-User Scheduling of Buffer-Aided Relay Systems

Pihe Hu, Cheng Li, Dingjie Xu +1

Multi-User scheduling is a challenging problem under the relaying scenarios. Traditional schemes, which are based on the instantaneous signal-to-interference-plus-noises ratios (SI…

cs.LG2023

Provably Efficient Iterated CVaR Reinforcement Learning with Function Approximation and Human Feedback

Yu Chen, Yihan Du, Pihe Hu +3

Risk-sensitive reinforcement learning (RL) aims to optimize policies that balance the expected reward and risk. In this paper, we present a novel risk-sensitive RL framework that e…