24 citations · 90 across the 37 of their papers we have counts for
43 papers
EStream: Fast and Memory-Efficient MoE Prefill through Expert Virtualization on Mobile NPUs
Junming Zhang, Zhenzhe Zheng, Fan Wu +2
Mobile vendors and application developers increasingly deploy LLMs on smartphones for diverse prefill-only services. Yet current systems rely mainly on dense models whose regular c…
An Efficient and Privacy-Preserving Architecture for Cross-Institutional Collaborative RAG
Chenxin Mao, Shangyu Liu, Zhenzhe Zheng +3
Retrieval-Augmented Generation (RAG) empowers LLMs with external knowledge, making cross-institutional domain-specific knowledge base integration a highly promising deployment para…
Optimizing Feature Extraction for On-device Model Inference with User Behavior Sequences
Chen Gong, Zhenzhe Zheng, Yiliu Chen +3
Machine learning models are widely integrated into modern mobile apps to analyze user behaviors and deliver personalized services. Ensuring low-latency on-device model execution is…
Guiding the Recommender: Information-Aware Auto-Bidding for Content Promotion
Yumou Liu, Zhenzhe Zheng, Jiang Rong +3
Modern content platforms offer paid promotion to mitigate cold start by allocating exposure via auctions. Our empirical analysis reveals a counterintuitive flaw in this paradigm: w…
Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs
Wei Zhou, Jun Zhou, Haoyu Wang +16
Data preparation aims to denoise raw datasets, uncover cross-dataset relationships, and extract valuable insights from them, which is essential for a wide range of data-centric app…
TTF: A Trapezoidal Temporal Fusion Framework for LTV Forecasting in Douyin
Yibing Wan, Zhengxiong Guan, Chaoli Zhang +5
In the user growth scenario, Internet companies invest heavily in paid acquisition channels to acquire new users. But sustainable growth depends on acquired users' generating lifet…