10 papers
Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving
Yifan Sui, Han Zhao, Rui Ma +6
LLM-powered agents execute tasks through a sequential loop of model generation and tool execution. Today's serving systems serialize this loop, leaving tool latency exposed on the…
World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
Weijie Wang, Xiaoxuan He, Youping Gu +9
Recent video foundation models demonstrate impressive visual synthesis but frequently suffer from geometric inconsistencies. While existing methods attempt to inject 3D priors via…
Accelerating Prefilling via Decoding-time Contribution Sparsity
Zhiyuan He, Yike Zhang, Chengruidong Zhang +3
Large Language Models (LLMs) incur quadratic attention complexity with input length, creating a major time bottleneck in the prefilling stage. Existing acceleration methods largely…
Joint Optimization of Handoff and Video Rate in LEO Satellite Networks
Kyoungjun Park, Zhiyuan He, Cheng Luo +4
Low Earth Orbit (LEO) satellite communication is a promising approach to providing Internet connectivity to users in many remote areas. As videos are likely to account for most tra…
Optimize Replica Server Placement in a Satellite Network
Zhiyuan He, Yi Xu, Cheng Luo +2
Satellite communication offers Internet connectivity to remote locations, such as villages, deserts, mountains, and at sea. However, transmitting content over satellite networks is…
VL Norm: Rethink Loss Aggregation in RLVR
Zhiyuan He, Xufang Luo, Yike Zhang +2
We propose VL Norm (Variance-reduced Length-dependent Normalization), a simple yet effective loss aggregation method tailored to the characteristic of dynamic generation lengths in…