3 papers
cs.CV2026
CoReDiT: Spatial Coherence-Guided Token Pruning and Reconstruction for Efficient Diffusion Transformers
Zhuojin Li, Hsin-Pai Cheng, Hong Cai +2
Diffusion Transformers (DiTs) deliver remarkable image and video generation quality but incur high computational cost, limiting scalability and on-device deployment. We introduce C…
cs.CV2026
A Study on Inference Latency for Vision Transformers on Mobile Devices
Zhuojin Li, Marco Paolieri, Leana Golubchik
Given the significant advances in machine learning techniques on mobile devices, particularly in the domain of computer vision, in this work we quantitatively study the performance…
cs.LG2026
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
Zhuojin Li, Marco Paolieri, Leana Golubchik
Deploying deep neural networks on mobile devices is increasingly important but remains challenging due to limited computing resources. On the other hand, their unified memory archi…