Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
CSR:Achieving 1 Bit Key-Value Cache via Sparse Representation
Hongxuan Zhang, Yao Zhao, Jiaqi Zheng +3
The emergence of long-context text applications utilizing large language models (LLMs) has presented significant scalability challenges, particularly in memory footprint. The linea…
cs.CL2024
Fast Chain-of-Thought: A Glance of Future from Parallel Decoding Leads to Answers Faster
Hongxuan Zhang, Zhining Liu, Yao Zhao +4
In this work, we propose FastCoT, a model-agnostic framework based on parallel decoding without any further training of an auxiliary model or modification to the LLM itself. FastCo…