5 papers
WhenLoss: Diagnosing Write and Retrieval Bottlenecks in Long-Context Memory Systems
Jiangnan Yu, Kisson Songqi Lin, Jilong Wu
Long-context memory systems often fail under fixed budgets, but end-to-end evaluation does not reveal whether evidence was discarded during compression or preserved but never retri…
STS: Efficient Sparse Attention with Speculative Token Sparsity
Ceyu Xu, Jiangnan Yu, Yongji Wu +1
The quadratic complexity of attention imposes severe memory and computational bottlenecks on Large Language Model (LLM) inference. This challenge is particularly acute for emerging…
DS-CIM: Digital Stochastic Computing-In-Memory Featuring Accurate OR-Accumulation via Sample Region Remapping for Edge AI Models
Kunming Shao, Liang Zhao, Jiangnan Yu +5
Stochastic computing (SC) offers hardware simplicity but suffers from low throughput, while high-throughput Digital Computing-in-Memory (DCIM) is bottlenecked by costly adder logic…
A Memory-Efficient Retrieval Architecture for RAG-Enabled Wearable Medical LLMs-Agents
Zhipeng Liao, Kunming Shao, Jiangnan Yu +5
With powerful and integrative large language models (LLMs), medical AI agents have demonstrated unique advantages in providing personalized medical consultations, continuous health…
DIRC-RAG: Accelerating Edge RAG with Robust High-Density and High-Loading-Bandwidth Digital In-ReRAM Computation
Kunming Shao, Zhipeng Liao, Jiangnan Yu +9
Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by integrating external knowledge retrieval but faces challenges on edge devices due to high storage, ene…