activity
20242026
collaborators

15 papers

cs.IR2026

GRACE: Generative Recommender Acceleration Engine for Real-Time Ads Retrieval

Zhou Fang, Yuhang Huang, Ang Zhang +12

Productionizing generative recommenders for high-volume, real-time ads retrieval creates two serving challenges: eligibility, ensuring that each generated ad is eligible for the re…

cs.IR2026

Diffusion-GR2: Diffusion Generative Reasoning Re-ranker

Zhuoxuan Zhang, Kangqi Ni, Yuhang Chen +12

Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-ordering a candidate list, but they are slow at inference: an autore…

cs.CL2026

Self-Guided Test-Time Training for Long-Context LLMs

Xinyu Zhu, Zhe Xu, Xiaohan Wei +10

Long-context processing has become increasingly important for large language models (LLMs), but simply extending the context window does not guarantee effective utilization of long…

cs.IR2026

SCOReD: Student-Aware CoT Optimization for Recommendation Distillation

Haz Sameen Shahgir, Yufei Li, Xiaohan Wei +8

Chain-of-thought (CoT) distillation in the recommendation domain is a necessary precursor to RL training, but raw teacher traces are ill-suited to this task. Large teachers approac…

cs.IR2026

GR2 Technical Report

Yufei Li, Zaiwei Zhang, Mingfu Liang +67

Industrial recommendation systems serve billions of users through a multi-stage funnel -- retrieval, early-stage ranking, and re-ranking -- where the final re-ranking step dispropo…

cs.IR2026

End-to-End Dynamic Sparsity for Resource-Adaptive LLM Inference

Yuhang Chen, Jinhao Duan, Ruichen Zhang +11

Large Language Models (LLMs) inference is typically deployed under a static resource assumption, where models execute a fixed computational graph regardless of the runtime environm…