7 papers
Learning to Seek Help: Dynamic Collaboration Between Small and Large Language Models
Hang Zeng, Xiangyu Liu, Yong Hu +5
Large language models (LLMs) offer strong capabilities but raise cost and privacy concerns, whereas small language models (SLMs) facilitate efficient and private local inference ye…
Optimizing Storage Overhead of User Behavior Log for ML-embedded Mobile Apps
Chen Gong, Yan Zhuang, Zhenzhe Zheng +4
Machine learning (ML) models are increasingly integrated into modern mobile apps to enable personalized and intelligent services. These models typically rely on rich input features…
CHORD: Customizing Hybrid-precision On-device Model for Sequential Recommendation with Device-cloud Collaboration
Tianqi Liu, Kairui Fu, Shengyu Zhang +5
With the advancement of mobile device capabilities, deploying reranking models directly on devices has become feasible, enabling real-time contextual recommendations. When migratin…
TSRec: Enhancing Repeat-Aware Recommendation from a Temporal-Sequential Perspective
Shigang Quan, Shui Liu, Zhenzhe Zheng +1
Repeat consumption, such as repurchasing items and relistening songs, is a common scenario in daily life. To model repeat consumption, the repeat-aware recommendation has been prop…
MERIT: A Merchant Incentive Ranking Model for Hotel Search & Ranking
Shigang Quan, Hailong Tan, Shui Liu +5
Online Travel Platforms (OTPs) have been working on improving their hotel Search & Ranking (S&R) systems that facilitate efficient matching between consumers and hotels. Existing O…
Pre: Enabling Deterministic Pushdown Automata for Faster Structured LLM Generation
Junyi Chen, Shihao Bai, Zaijun Wang +7
Extensive LLM applications demand efficient structured generations, particularly for LR(1) grammars, to produce outputs in specified formats (e.g., JSON). Existing methods primaril…