6 papers
STEP: Warm-Started Visuomotor Policies with Spatiotemporal Consistency Prediction
Jinhao Li, Yuxuan Cong, Yingqiao Wang +5
Diffusion policies have recently emerged as a powerful paradigm for visuomotor control in robotic manipulation due to their ability to model the distribution of action sequences an…
Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective
Jinhao Li, Jiaming Xu, Shan Huang +9
Large Language Models (LLMs) have demonstrated remarkable capabilities across various fields, from natural language understanding to text generation. Compared to non-generative LLM…
DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale
Ziyang Zheng, Shan Huang, Jianyuan Zhong +4
Circuit representation learning has become pivotal in electronic design automation, enabling critical tasks such as testability analysis, logic reasoning, power estimation, and SAT…
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
Jiaming Xu, Jiayi Pan, Yongkang Zhou +5
Early exiting has recently emerged as a promising technique for accelerating large language models (LLMs) by effectively reducing the hardware computation and memory access. In thi…
Automating Energy-Efficient GPU Kernel Generation: A Fast Search-Based Compilation Approach
Yijia Zhang, Zhihong Gou, Shijie Cao +4
Deep Neural Networks (DNNs) have revolutionized various fields, but their deployment on GPUs often leads to significant energy consumption. Unlike existing methods for reducing GPU…
SoftmAP: Software-Hardware Co-design for Integer-Only Softmax on Associative Processors
Mariam Rakka, Jinhao Li, Guohao Dai +3
Recent research efforts focus on reducing the computational and memory overheads of Large Language Models (LLMs) to make them feasible on resource-constrained devices. Despite adva…