activity
20242026
collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2025

When Long Helps Short: How Context Length in Supervised Fine-tuning Affects Behavior of Large Language Models

Yingming Zheng, Hanqi Li, Kai Yu +1

Large language models (LLMs) have achieved impressive performance across natural language processing (NLP) tasks. As real-world applications increasingly demand longer context wind…

cs.CL2025

Compressing KV Cache for Long-Context LLM Inference with Inter-Layer Attention Similarity

Da Ma, Lu Chen, Situo Zhang +8

The rapid expansion of context window sizes in Large Language Models~(LLMs) has enabled them to tackle increasingly complex tasks involving lengthy documents. However, this progres…

cs.CL2025

NeuSym-RAG: Hybrid Neural Symbolic Retrieval with Multiview Structuring for PDF Question Answering

Ruisheng Cao, Hanchong Zhang, Tiancheng Huang +8

The increasing number of academic papers poses significant challenges for researchers to efficiently acquire key details. While retrieval augmented generation (RAG) shows great pro…

cs.CL2024

Evolving Subnetwork Training for Large Language Models

Hanqi Li, Lu Chen, Da Ma +3

Large language models have ushered in a new era of artificial intelligence research. However, their substantial training costs hinder further development and widespread adoption. I…

cs.CL2024

Sparsity-Accelerated Training for Large Language Models

Da Ma, Lu Chen, Pengyu Wang +6

Large language models (LLMs) have demonstrated proficiency across various natural language processing (NLP) tasks but often require additional training, such as continual pre-train…