5 papers · 1 filter
ProDVI: Programmatic Dynamics Priors for Value Network Initialization
Xinwei Liu, Junyuan Liang, Jianting Zhang +1
Deep Reinforcement Learning (RL) is notoriously sample inefficient. One contributing factor is that RL agents are typically initialized from scratch, forcing them to acquire task-r…
Observation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous Control
Xinwei Liu, Junyuan Liang, Jianting Zhang +1
Sample-efficient policy learning from pixels is a long-standing challenge in reinforcement learning (RL). Recent dynamics-based representation learning methods have significantly i…
NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning
Xinwei Liu, Junyuan Liang, Zicong Hong +2
Augmenting model-free reinforcement learning (RL) with representations learned through observation dynamics prediction (observation-predictive RL) can improve sample efficiency and…
DyMoE: Dynamic Expert Orchestration with Mixed-Precision Quantization for Efficient MoE Inference on Edge
Yuegui Huang, Zhiyuan Fang, Weiqi Luo +3
Despite the computational efficiency of MoE models, the excessive memory footprint and I/O overhead inherent in multi-expert architectures pose formidable challenges for real-time…
Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline
Zhiyuan Fang, Yuegui Huang, Zicong Hong +5
Mixture of Experts (MoE), with its distinctive sparse structure, enables the scaling of language models up to trillions of parameters without significantly increasing computational…