5 papers
Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation
Shuaiyi Li, Zhisong Zhang, Yan Wang +5
Block attention, which processes the input as separate blocks that cannot attend to one another, offers significant potential to improve KV cache reuse in long-context scenarios su…
PEARL: Towards Permutation-Resilient LLMs
Liang Chen, Li Shen, Yang Deng +3
The in-context learning (ICL) capability of large language models (LLMs) enables them to perform challenging tasks using provided demonstrations. However, ICL is highly sensitive t…
Consecutive Batch Model Editing with HooK Layers
Shuaiyi Li, Yang Deng, Deng Cai +3
As the typical retraining paradigm is unacceptably time- and resource-consuming, researchers are turning to model editing to find an effective way that supports both consecutive an…
UniMS-RAG: A Unified Multi-source Retrieval-Augmented Generation for Personalized Dialogue Systems
Hongru Wang, Wenyu Huang, Yang Deng +6
Large Language Models (LLMs) has shown exceptional capabilities in many natual language understanding and generation tasks. However, the personalization issue still remains a much-…
WatME: Towards Lossless Watermarking Through Lexical Redundancy
Liang Chen, Yatao Bian, Yang Deng +4
Text watermarking has emerged as a pivotal technique for identifying machine-generated text. However, existing methods often rely on arbitrary vocabulary partitioning during decodi…