2 papers
cs.LG2026
CSAttention: Centroid-Scoring Attention for Accelerating LLM Inference
Chuxu Song, Zhencan Peng, Jiuqi Wei +1
Long-context LLMs increasingly rely on extended, reusable prefill prompts for agents and domain Q&A, pushing attention and KV-cache to become the dominant decode-time bottlenecks.…
cs.DB2025
Near-Duplicate Text Alignment under Weighted Jaccard Similarity
Yuheng Zhang, Miao Qiao, Zhencan Peng +1
Near-duplicate text alignment is the task of identifying, among the texts in a corpus, all the subsequences (substrings) that are similar to a given query. Traditional approaches r…