3 papers
cs.LG2026
ASTRA: Communication-Efficient Acceleration for Multi-Device Transformer Inference
Xiao Liu, Lijun Zhang, Deepak Ganesan +1
Multi-device inference can reduce Transformer latency by parallelizing computation. However, existing methods require high inter-device bandwidth, making them impractical for bandw…
cs.CV2026
Aligned Vector Quantization for Edge-Cloud Collabrative Vision-Language Models
Xiao Liu, Lijun Zhang, Deepak Ganesan +1
Vision Language Models (VLMs) are central to Visual Question Answering (VQA) systems and are typically deployed in the cloud due to their high computational demands. However, this…
cs.NE2025
Adaptive Spiking with Plasticity for Energy Aware Neuromorphic Systems
Eduardo Calle-Ortiz, Hui Guan, Deepak Ganesan +1
This paper presents ASPEN, a novel energy-aware technique for neuromorphic systems that could unleash the future of intelligent, always-on, ultra-low-power, and low-burden wearable…