3 papers
cs.CL2025
EvoP: Robust LLM Inference via Evolutionary Pruning
Shangyu Wu, Hongchao Du, Ying Xiong +4
Large Language Models (LLMs) have achieved remarkable success in natural language processing tasks, but their massive size and computational demands hinder their deployment in reso…
cs.CL2025
AATS: Retrieval-Based KV Cache Reduction via Windowed Rotary Position Embedding and Query-Aware Vector Quantization
Junhui He, Junna Xing, Nan Wang +6
Long context large language models (LLMs) pose significant challenges for efficient serving due to the large memory footprint and high access overhead of KV cache. Retrieval-based…
cs.CV2024
GeneQuery: A General QA-based Framework for Spatial Gene Expression Predictions from Histology Images
Ying Xiong, Linjing Liu, Yufei Cui +4
Gene expression profiling provides profound insights into molecular mechanisms, but its time-consuming and costly nature often presents significant challenges. In contrast, whole-s…