activity
20242026
most citedSmartMem: Layout Transformation Elimination and Adaptation for Efficient DNN Execution on Mobile

11 citations · 11 across the 5 of their papers we have counts for

collaborators

7 papers

cs.LG2026

GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents

Zhe Ren, Yibo Yang, Yimeng Chen +7

Memory benchmarks for LLM agents largely assume single-user settings, leaving shared assistants for hospitals, workplaces, campuses, and households understudied. In these deploymen…

cs.DC2026

FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations

Zhihao Shu, Md Musfiqur Rahman Sanim, Hangyu Zheng +4

The increasing size and complexity of modern deep neural networks (DNNs) pose significant challenges for on-device inference on mobile GPUs, with limited memory and computational r…

cs.CV2025

Optimizing 3D Gaussian Splattering for Mobile GPUs

Md Musfiqur Rahman Sanim, Zhihao Shu, Bahram Afsharmanesh +5

Image-based 3D scene reconstruction, which transforms multi-view images into a structured 3D representation of the surrounding environment, is a common task across many modern appl…

cs.LG2024

LazyDiT: Lazy Learning for the Acceleration of Diffusion Transformers

Xuan Shen, Zhao Song, Yufa Zhou +12

Diffusion Transformers have emerged as the preeminent models for a wide array of generative tasks, demonstrating superior performance and efficacy across various applications. The…

cs.CV2024

Data Overfitting for On-Device Super-Resolution with Dynamic Algorithm and Compiler Co-Design

Gen Li, Zhihao Shu, Jie Ji +4

Deep neural networks (DNNs) are frequently employed in a variety of computer vision applications. Nowadays, an emerging trend in the current video distribution system is to take ad…

cs.LG202411 cited

SmartMem: Layout Transformation Elimination and Adaptation for Efficient DNN Execution on Mobile

Wei Niu, Md Musfiqur Rahman Sanim, Zhihao Shu +5

This work is motivated by recent developments in Deep Neural Networks, particularly the Transformer architectures underlying applications such as ChatGPT, and the need for performi…