2 papers
cs.DC2026
Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving
Yifan Sui, Han Zhao, Rui Ma +6
LLM-powered agents execute tasks through a sequential loop of model generation and tool execution. Today's serving systems serialize this loop, leaving tool latency exposed on the…
cs.IR2026
Efficient Personalized Reranking with Semi-Autoregressive Generation and Online Knowledge Distillation
Kai Cheng, Hao Wang, Wei Guo +4
Generative models offer a promising paradigm for the final stage reranking in multi-stage recommender systems, with the ability to capture inter-item dependencies within reranked l…