activity
20242026
collaborators

7 papers

cs.LG2026

Optimizing Feature Extraction for On-device Model Inference with User Behavior Sequences

Chen Gong, Zhenzhe Zheng, Yiliu Chen +3

Machine learning models are widely integrated into modern mobile apps to analyze user behaviors and deliver personalized services. Ensuring low-latency on-device model execution is…

cs.GT2026

Guiding the Recommender: Information-Aware Auto-Bidding for Content Promotion

Yumou Liu, Zhenzhe Zheng, Jiang Rong +3

Modern content platforms offer paid promotion to mitigate cold start by allocating exposure via auctions. Our empirical analysis reveals a counterintuitive flaw in this paradigm: w…

cs.LG2025

TTF: A Trapezoidal Temporal Fusion Framework for LTV Forecasting in Douyin

Yibing Wan, Zhengxiong Guan, Chaoli Zhang +5

In the user growth scenario, Internet companies invest heavily in paid acquisition channels to acquire new users. But sustainable growth depends on acquired users' generating lifet…

cs.LG2025

A Two-Stage Data Selection Framework for Data-Efficient Model Training on Edge Devices

Chen Gong, Rui Xing, Zhenzhe Zheng +1

The demand for machine learning (ML) model training on edge devices is escalating due to data privacy and personalized service needs. However, we observe that current on-device mod…

cs.LG2025

Efficient Distributed Retrieval-Augmented Generation for Enhancing Language Model Performance

Shangyu Liu, Zhenzhe Zheng, Xiaoyao Huang +3

Small language models (SLMs) support efficient deployments on resource-constrained edge devices, but their limited capacity compromises inference performance. Retrieval-augmented g…

cs.CL2025

AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference

Zhuomin He, Yizhen Yao, Pengfei Zuo +4

Long-context large language models (LLMs) inference is increasingly critical, motivating a number of studies devoted to alleviating the substantial storage and computational costs…