5 papers
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
Yinpeng Wu, Yitong Chen, Lixiang Wang +3
Device-side Large Language Models (LLMs) have grown explosively, offering stronger privacy and higher availability than their cloud-side counterparts. During LLM inference, both th…
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
Yinpeng Wu, Yitong Chen, Lixiang Wang +3
Device-side Large Language Models (LLMs) have witnessed explosive growth, offering higher privacy and availability compared to cloud-side LLMs. During LLM inference, both model wei…
CS3: Efficient Online Capability Synergy for Two-Tower Recommendation
Lixiang Wang, Shaoyun Shi, Peng Wang +2
To balance effectiveness and efficiency in recommender systems, multi-stage pipelines employ lightweight two-tower models for large-scale candidate retrieval. However, their isolat…
CS3: Efficient Online Capability Synergy for Two-Tower Recommendation
Lixiang Wang, Shaoyun Shi, Peng Wang +2
To balance effectiveness and efficiency in recommender systems, multi-stage pipelines commonly use lightweight two-tower models for large-scale candidate retrieval. However, the is…
Generative Recommendation for Large-Scale Advertising
Ben Xue, Dan Liu, Lixiang Wang +27
Generative recommendation has recently attracted widespread attention in industry due to its potential for scaling and stronger model capacity. However, deploying real-time generat…