works on

From the 1 of 17 linked papers with an AI index.

activity
20242026
collaborators

17 papers

cs.RO2026

RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy

Zhengyang Yan, Junhao Li, Fangqi Zhu +6

RedFlow is an offline reinforcement learning framework that turns failure experiences into action-level corrective supervision for flow-matching vision‑language‑action policies, im…

cs.LG2026

NASDAQ: Normalized Observation Space Dynamics-Augmented Q-Learning

Xinwei Liu, Junyuan Liang, Zicong Hong +2

Augmenting model-free reinforcement learning (RL) with representations learned through observation dynamics prediction (observation-predictive RL) can improve sample efficiency and…

cs.LG2026

MosaicQuant: Inlier-Outlier Disaggregation for Unified 4-Bit LLM Quantization

Yangjia Hu, Haodong Wang, Zicong Hong +8

4-bit quantization significantly reduces the memory footprint and accelerates the inference of large language models (LLMs). However, its limited bit-width representation struggles…

cs.LG2026

Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement

Qianli Liu, Kaibin Guo, Zicong Hong +5

Expert parallelism has become the prevailing paradigm to serve Mixture-of-Experts (MoE) models. Its efficiency depends on the communication and computation latencies of the GPUs, w…

cs.DC2026

TwinQuant: Learnable Subspace Decomposition for 4-Bit LLM Quantization

Haodong Wang, Junjie Liu, Zicong Hong +4

4-bit quantization reduces the memory footprint and latency of large language model inference, but its aggressive precision reduction can severely degrade accuracy. Prior methods a…

cs.CL2026

KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference

Jian Lin, Jiazhi Mi, Zicong Hong +5

Supporting long-context LLMs is challenging due to the substantial memory demands of the key-value (KV) cache. Existing offloading systems store the full cache in host memory and s…