most citedApt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving

5 citations · 9 across the 4 of their papers we have counts for

collaborators

5 papers

cs.DB2025

RelServe: Fast LLM Inference Serving on Relational Data

Xin Zhang, Shihong Gao, Yanyan Shen +2

The use of Large Language Models (LLMs) for querying relational data has given rise to relQuery, a workload pattern that applies templated LLM calls to structured tables. As relQue…

q-fin.ST2025

Momentum-integrated Multi-task Stock Recommendation with Converge-based Optimization

Hao Wang, Jingshu Peng, Yanyan Shen +4

Stock recommendation is critical in Fintech applications, which leverage price series and alternative information to estimate future stock performance. Traditional time-series fore…

cs.IR20254 cited

A Framework for Elastic Adaptation of User Multiple Intents in Sequential Recommendation

Zhikai Wang, Yanyan Shen

Recently, substantial research has been conducted on sequential recommendation, with the objective of forecasting the subsequent item by leveraging a user's historical sequence of…

cs.LG20255 cited

Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving

Shihong Gao, Xin Zhang, Yanyan Shen +1

Large language model (LLM) inference serving systems are essential to various LLM-based applications. As demand for LLM services continues to grow, scaling these systems to handle…

cs.DC2025

Improving the End-to-End Efficiency of Offline Inference for Multi-LLM Applications Based on Sampling and Simulation

Jingzhi Fang, Yanyan Shen, Yue Wang +1

As large language models (LLMs) have shown great success in many tasks, they are used in various applications. While a lot of works have focused on the efficiency of single-LLM app…