3 papers
cs.IR2026
CS3: Efficient Online Capability Synergy for Two-Tower Recommendation
Lixiang Wang, Shaoyun Shi, Peng Wang +2
To balance effectiveness and efficiency in recommender systems, multi-stage pipelines employ lightweight two-tower models for large-scale candidate retrieval. However, their isolat…
cs.IR2026
CS3: Efficient Online Capability Synergy for Two-Tower Recommendation
Lixiang Wang, Shaoyun Shi, Peng Wang +2
To balance effectiveness and efficiency in recommender systems, multi-stage pipelines commonly use lightweight two-tower models for large-scale candidate retrieval. However, the is…
cs.CL2026
HyLRA: Hybrid Layer Reuse Attention for Efficient Long-Context Inference
Xuan Ai, Qingqing Yang, Peng Wang +4
Long-context inference in Large Language Models (LLMs) is bottlenecked by the quadratic computation complexity of attention and the substantial memory footprint of Key-Value (KV) c…