8 papers
Agent-Assisted Side-Channel Attacks on Non-Prefix KV Cache in RAG
He Sun, Shinan Liu, Siyuan Ma +3
Modern Large Language Model (LLM) serving engines increasingly rely on Retrieval-Augmented Generation (RAG) and non-prefix Key-Value (KV) cache fusion to accelerate long-context, m…
MACReD: A Multi-Agent Collaborative Reasoning Framework for Reaction Diagram Parsing
Chuang Tang, Chenhao Lin, Yin Xu +5
Parsing chemical reaction diagrams from scientific literature is challenging due to heterogeneous layouts, intertwined visual elements, and the difficulty of integrating recognitio…
HillInfer: Efficient Long-Context LLM Inference on the Edge with Hierarchical KV Eviction using SmartSSD
He Sun, Shinan Liu, Li Li +1
Deploying Large Language Models (LLMs) on memory-constrained AI Personal Computers (AIPCs) enables low-latency, privacy-preserving inference, but long-context generation is fundame…
CooperLLM: Cloud-Edge-End Cooperative Federated Fine-tuning for LLMs via ZOO-based Gradient Correction
He Sun, Jinrui Zhou, Li Li +1
Large Language Models (LLMs) perform well on many NLP tasks, but fine-tuning them on resource-constrained mobile devices is challenging due to high memory and computation costs, de…
Enhancing Automated Paper Reproduction via Prompt-Free Collaborative Agents
Zijie Lin, Qilin Cai, Liang Shen +1
Automated paper reproduction has emerged as a promising approach to accelerate scientific research, employing multi-step workflow frameworks to systematically convert academic pape…
SFedKD: Sequential Federated Learning with Discrepancy-Aware Multi-Teacher Knowledge Distillation
Haotian Xu, Jinrui Zhou, Xichong Zhang +3
Federated Learning (FL) is a distributed machine learning paradigm which coordinates multiple clients to collaboratively train a global model via a central server. Sequential Feder…