5 papers
HEMERA: A Heterogeneous Memory-Centric Accelerator with Recursive Dataflow for Edge-Constrained State-Space-Duality Models Inference
Hao Ding, Ling Liang, Ruitong Qiao +9
Structured State Space Models (SSMs), such as Mamba, enable efficient long-sequence modeling with linear time complexity. Recent implementations realize this capability through Str…
DynaGraph: Lightweight Multi-Model Interaction Framework via Dynamic Topological Reconfiguration
Yanxing Guo, Zihao Zheng, Fangzhou Wu +4
Tackling complex reasoning tasks typically relies on massive monolithic LLMs, which suffer from severe computational redundancy. While task decomposition through structured pipelin…
AdaMemento: Adaptive Memory-Assisted Policy Optimization for Reinforcement Learning
Renye Yan, Yaozhong Gan, You Wu +4
In sparse reward scenarios of reinforcement learning (RL), the memory mechanism provides promising shortcuts to policy optimization by reflecting on past experiences like humans. H…
NASiC: 3D NAND-based CAM-Selected Multibit CIM Architecture for Efficient On-Device Mixture-of-Experts LLM Inference
Weikai Xu, Meng Li, Shuzhang Zhong +7
The Mixture-of-Experts (MoE) models have emerged as the state-of-the-art paradigm for scaling up large language models (LLMs) without proportionally increased computational cost. H…
Inner-Probe: Discovering Copyright-related Data Generation in LLM Architecture
Qichao Ma, Rui-Jie Zhu, Peiye Liu +8
Large Language Models (LLMs) utilize extensive knowledge databases and show powerful text generation ability. However, their reliance on high-quality copyrighted datasets raises co…