works on

From the 1 of 14 linked papers with an AI index.

most citedWisdom of Committee: Diverse Distillation from Large Foundation Models and Domain Experts

1 citations · 1 across the 6 of their papers we have counts for

collaborators

14 papers

cs.IR2026

Tokens are All You Need: Dual-purpose Semantic IDs for Achieving LLM-Level I/O Efficiency in recommendation systems

Baolei Li, Yiping Yuan, Yilin Zheng +6

Large-scale recommendation systems face "Memory Wall" bottlenecks due to massive, dense embedding tables. While generative retrieval uses discrete tokens for IDs, high-dimensional…

cs.IR2026

LLM-Based User Personas for Recommendations at Scale

Haoting Wang, Haokai Lu, Zheyun Feng +14

The paper presents a framework that uses large language models to generate natural-language user interest personas in real time for a large‑scale video recommendation system, emplo…

cs.IR2026

TokenMinds: Pretrained User Tokens and Embeddings for User Understanding in Large Recommender Systems

Qingyun Liu, Bo Yan, Yang Liu +15

User modeling in industrial recommender systems typically produces dense embeddings, which suffer from representational constraints inherent to fixed-dimensional vectors. An emergi…

cs.IR2026

Token Factory: Efficiently Integrating Diverse Signals into Large Recommendation Models

Xilun Chen, Shao-Chuan Wang, Baykal Cakici +6

Large Recommendation Models (LRMs) have demonstrated promising capabilities in industry-scale recommendation tasks. However, holistically integrating traditional signals into these…

cs.LG20261 cited

Wisdom of Committee: Diverse Distillation from Large Foundation Models and Domain Experts

Zichang Liu, Qingyun Liu, Yuening Li +6

Knowledge distillation from foundation models to compact domain models is challenging due to substantial gaps in capacity, architecture, and modality. For example, in our experimen…

cs.CL2026

ORBIT: Preserving Foundational Language Capabilities in GenRetrieval via Origin-Regulated Merging

Neha Verma, Nikhil Mehta, Shao-Chuan Wang +7

Despite the rapid advancements in large language model (LLM) development, fine-tuning them for specific tasks often results in the catastrophic forgetting of their general, languag…