collaborators

5 papers

cs.AR2026

Potential Applications of HBF in LLM Serving Systems

Yihan Yin, Yinlun Zhao, Zhixin Yun +8

LLM serving is increasingly constrained by memory capacity as model weights, KV caches, and the number of served model variants continue to grow. This report examines High-Bandwidt…

cond-mat.mtrl-sci2026

High-speed and high-gain graphene photovoltaic phototransistor gated by a van der Waals heterojunction

Yihan Yin, Jiayi Zhang, Xiaolong Zhang +5

Two-dimensional (2D) material-based phototransistors offer a unique combination of optical sensing, signal amplification, and logic operation within a single device, yet fundamenta…

cs.AR2026

CODA: Algorithm-Hardware Co-design for Edge Video Diffusion via NMP-Enabled Compute-Cache Operator Disaggregation

Yuanpeng Zhang, YuXuan Wu, Yitong Xiao +6

Deploying Video Diffusion Models (VDMs) on edge devices is appealing for localized and privacy-preserving generation, but their iterative Transformer-based denoising remains too sl…

cs.AR2026

AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator

Chenhao Xue, Yukun Wang, An Guo +11

SRAM-based compute-in-memory (CIM) offers high computational density and energy efficiency for deep neural network (DNN) accelerators, but its limited capacity causes on/off-chip d…

cs.AR2026

Hardware-Software Co-design for 3D-DRAM-based LLM Serving Accelerator

Cong Li, Yihan Yin, Chenhao Xue +7

Large language models (LLMs) have been widely deployed for online generative services, where numerous LLM instances jointly handle workloads with fluctuating request arrival rates…