3 papers
cs.CL2025
Put Teacher in Student's Shoes: Cross-Distillation for Ultra-compact Model Compression Framework
Maolin Wang, Jun Chu, Sicong Xie +4
In the era of mobile computing, deploying efficient Natural Language Processing (NLP) models in resource-restricted edge settings presents significant challenges, particularly in e…
cs.IR2024
DNS-Rec: Data-aware Neural Architecture Search for Recommender Systems
Sheng Zhang, Maolin Wang, Yao Zhao +6
In the era of data proliferation, efficiently sifting through vast information to extract meaningful insights has become increasingly crucial. This paper addresses the computationa…
cs.CL2024
CSR:Achieving 1 Bit Key-Value Cache via Sparse Representation
Hongxuan Zhang, Yao Zhao, Jiaqi Zheng +3
The emergence of long-context text applications utilizing large language models (LLMs) has presented significant scalability challenges, particularly in memory footprint. The linea…