4 papers · 1 filter
CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving
Dong Liu, Yanxuan Yu
Large Language Models (LLMs) have revolutionized natural language processing tasks, but their deployment in datacenter environments faces significant challenges due to the massive…
Cognitive Load Traces as Symbolic and Visual Accounts of Deep Model Cognition
Dong Liu, Yanxuan Yu
We propose \textbf{Cognitive Load Traces} (CLTs) as a mid-level interpretability framework for deep models, inspired by Cognitive Load Theory in human cognition. CLTs are defined a…
HSGM: Hierarchical Segment-Graph Memory for Scalable Long-Text Semantics
Dong Liu, Yanxuan Yu
Semantic parsing of long documents remains challenging due to quadratic growth in pairwise composition and memory requirements. We introduce \textbf{Hierarchical Segment-Graph Memo…
QuickMerge++: Fast Token Merging with Autoregressive Prior
Dong Liu, Yanxuan Yu
As generative models scale to larger inputs across language, vision, and video domains, the cost of token-level computation has become a key bottleneck. While prior work suggests t…