works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.LG2026

ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation

Yuxin Chen, Liang Luo, Buyun Zhang +44

The paper introduces ROCS, a request-oriented compute sharing framework that restructures recommendation inference to evaluate shared request features once per request rather than…

cs.LG2026

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale

Liang Luo, Yinbin Ma, Quanyu Zhu +21

Recent GPU generations deliver significantly higher FLOPs using lower-precision arithmetic, such as FP8. While successfully applied to large language models (LLMs), its adoption in…

cs.IR2026

ReasonRec: A Reasoning-Augmented Multimodal Agent for Unified Recommendation

Yihua Zhang, Mingfu Liang, Jiyan Yang +11

Recent advances in multimodal recommenders excel at feature fusion but remain opaque and inefficient decision-makers, lacking explicit reasoning and self-awareness of uncertainty.…

cs.IR2026

The Efficiency vs. Accuracy Trade-off: Optimizing RAG-Enhanced LLM Recommender Systems Using Multi-Head Early Exit

Huixue Zhou, Hengrui Gu, Xi Liu +15

The deployment of Large Language Models (LLMs) in recommender systems for predicting Click-Through Rates (CTR) necessitates a delicate balance between computational efficiency and…

cs.LG2026

Implicit Turn-Wise Policy Optimization for Proactive User-LLM Interaction

Haoyu Wang, Yuxin Chen, Liang Luo +3

Multi-turn human-AI collaboration is fundamental to deploying interactive services such as adaptive tutoring, conversational recommendation, and professional consultation. However,…

cs.IR2025

Meta Lattice: Model Space Redesign for Cost-Effective Industry-Scale Ads Recommendations

Liang Luo, Yuxin Chen, Zhengyu Zhang +39

The rapidly evolving landscape of products, surfaces, policies, and regulations poses significant challenges for deploying state-of-the-art recommendation models at industry scale,…