collaborators

8 papers

cs.LG2026

ProDVI: Programmatic Dynamics Priors for Value Network Initialization

Xinwei Liu, Junyuan Liang, Jianting Zhang +1

Deep Reinforcement Learning (RL) is notoriously sample inefficient. One contributing factor is that RL agents are typically initialized from scratch, forcing them to acquire task-r…

cs.LG2026

Observation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous Control

Xinwei Liu, Junyuan Liang, Jianting Zhang +1

Sample-efficient policy learning from pixels is a long-standing challenge in reinforcement learning (RL). Recent dynamics-based representation learning methods have significantly i…

cs.AI2026

Hawk: Harnessing Hardware-Aware Knowledge for High-Performance NPU Kernel Generation

Junyi Wen, Ruiyan Zhuang, Yongjia Xu +7

Developing high-performance kernels for Neural Processing Units (NPUs) is a critical industry bottleneck, requiring developers to manually navigate implicit hardware constraints an…

cs.LG2026

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning

Xinwei Liu, Junyuan Liang, Zicong Hong +2

Augmenting model-free reinforcement learning (RL) with representations learned through observation dynamics prediction (observation-predictive RL) can improve sample efficiency and…

cs.LG2026

DyMoE: Dynamic Expert Orchestration with Mixed-Precision Quantization for Efficient MoE Inference on Edge

Yuegui Huang, Zhiyuan Fang, Weiqi Luo +3

Despite the computational efficiency of MoE models, the excessive memory footprint and I/O overhead inherent in multi-expert architectures pose formidable challenges for real-time…

cs.CL2025

Krul: Efficient State Restoration for Multi-turn Conversations with Dynamic Cross-layer KV Sharing

Junyi Wen, Junyuan Liang, Zicong Hong +3

Efficient state restoration in multi-turn conversations with large language models (LLMs) remains a critical challenge, primarily due to the overhead of recomputing or loading full…