3 papers
cs.LG2026
Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-Training
Hanfeng Lu, Tianyu Feng, Suyi Li +8
Vision-language models (VLMs) enable embodied agents to reason and act from visual observations and language instructions. Reinforcement learning (RL) post-training enhances these…
cs.AR2026
MoE Expert Execution in Disaggregated LLM Serving with a High-Bandwidth ReRAM Near-Memory Architecture
Kunming Shao, Ming Zeng, Xin Yuan +5
Attention-FFN disaggregation maps LLM modules to specialized pools, creating an opening to keep Mixture-of-Experts (MoE) weights resident in a high-bandwidth FFN pool. Decode SLOs,…
cs.CR2026
CipherSight: Robust Website Fingerprinting via Record-Resource Semantic Supervision under Distribution Shifts
Runhan Song, Qiqi Liu, Chuanzhou Pan +6
HTTPS website fingerprinting (WF) aims to identify visited websites from metadata observable in encrypted traffic. However, real-world deployments introduce a significant out-of-di…