11 citations · 11 across the 5 of their papers we have counts for
6 papers
Mask-Aware Execution for Efficient JEPA Training
Md Musfiqur Rahman Sanim, Zhihao Shu, Bahram Afsharmanesh +3
Joint Embedding Predictive Architectures (JEPAs) are becoming a core representation-learning primitive and a building block for latent world models across vision, video, audio, bra…
TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching
Zhihao Shu, Md Musfiqur Rahman Sanim, Jie Hu +4
Large language models (LLMs) are moving onto mobile devices for increasingly diverse workloads over text, images, video, and audio. These applications often require long contexts,…
FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations
Zhihao Shu, Md Musfiqur Rahman Sanim, Hangyu Zheng +4
The increasing size and complexity of modern deep neural networks (DNNs) pose significant challenges for on-device inference on mobile GPUs, with limited memory and computational r…
Optimizing 3D Gaussian Splattering for Mobile GPUs
Md Musfiqur Rahman Sanim, Zhihao Shu, Bahram Afsharmanesh +5
Image-based 3D scene reconstruction, which transforms multi-view images into a structured 3D representation of the surrounding environment, is a common task across many modern appl…
SmartMem: Layout Transformation Elimination and Adaptation for Efficient DNN Execution on Mobile
Wei Niu, Md Musfiqur Rahman Sanim, Zhihao Shu +5
This work is motivated by recent developments in Deep Neural Networks, particularly the Transformer architectures underlying applications such as ChatGPT, and the need for performi…
A Simple 2-Approximation Algorithm For Minimum Manhattan Network Problem
Md. Musfiqur Rahman Sanim, Safrunnesa Saira, Fatin Faiaz Ahsan +2
Given a n points in two dimensional space, a Manhattan Network G is a network that connects all n points with either horizontal or vertical edges, with the property that for any tw…