2 papers
cs.CL2024
S2-Attention: Hardware-Aware Context Sharding Among Attention Heads
Xihui Lin, Yunan Zhang, Suyu Ge +5
Sparse attention, which selectively attends to a subset of tokens in the context was supposed to be efficient. However, its theoretical reduction in FLOPs has rarely translated int…
cs.IR2024
GenSERP: Large Language Models for Whole Page Presentation
Zhenning Zhang, Yunan Zhang, Suyu Ge +4
The advent of large language models (LLMs) brings an opportunity to minimize the effort in search engine result page (SERP) organization. In this paper, we propose GenSERP, a frame…