2 papers
cs.AI2026
MemoSight: Unifying Context Compression and Multi Token Prediction for Reasoning Acceleration
Xinyu Liu, Xin Liu, Bo Jin +8
While chain-of-thought (CoT) reasoning enables LLMs to solve challenging reasoning tasks, the linear growth of the KV cache leads to substantial memory and inference overhead. Exis…
cs.LG2025
FinLoRA: Finetuning Quantized Financial Large Language Models Using Low-Rank Adaptation
Dannong Wang, Daniel Kim, Bo Jin +4
Finetuned large language models (LLMs) have shown remarkable performance in financial tasks, such as sentiment analysis and information retrieval. Due to privacy concerns, finetuni…