7 papers
CFSP: An Efficient Structured Pruning Framework for LLMs with Coarse-to-Fine Activation Information
Yuxin Wang, Minghua Ma, Zekun Wang +7
The colossal parameters and computational overhead of Large Language Models (LLMs) challenge their real-world applications. Network pruning, which targets unstructured or structure…
Meaningful Learning: Enhancing Abstract Reasoning in Large Language Models via Generic Fact Guidance
Kai Xiong, Xiao Ding, Ting Liu +5
Large language models (LLMs) have developed impressive performance and strong explainability across various reasoning scenarios, marking a significant stride towards mimicking huma…
How does Architecture Influence the Base Capabilities of Pre-trained Language Models? A Case Study Based on FFN-Wider and MoE Transformers
Xin Lu, Yanyan Zhao, Bing Qin +3
Pre-trained language models have been proven to possess strong base capabilities, which not only excel in in-distribution language modeling but also show powerful abilities in out-…
Advancing Large Language Model Attribution through Self-Improving
Lei Huang, Xiaocheng Feng, Weitao Ma +7
Teaching large language models (LLMs) to generate text with citations to evidence sources can mitigate hallucinations and enhance verifiability in information-seeking systems. Howe…
Length Extrapolation of Transformers: A Survey from the Perspective of Positional Encoding
Liang Zhao, Xiachong Feng, Xiaocheng Feng +6
Built upon the Transformer, large language models (LLMs) have captured worldwide attention due to their remarkable abilities. Nevertheless, all Transformer-based models including L…
GlobeSumm: A Challenging Benchmark Towards Unifying Multi-lingual, Cross-lingual and Multi-document News Summarization
Yangfan Ye, Xiachong Feng, Xiaocheng Feng +6
News summarization in today's global scene can be daunting with its flood of multilingual content and varied viewpoints from different sources. However, current studies often negle…