Publications (8)
FlashFPS: Efficient Farthest Point Sampling for Large-Scale Point Clouds via Pruning and Caching
Yuzhe Fu, Hancheng Ye, Cong Guo +7
Point-based Neural Networks (PNNs) have become a key approach for point cloud processing. However, a core operation in these models, Farthest Point Sampling (FPS), often introduces…
A Survey: Collaborative Hardware and Software Design in the Era of Large Language Models
Cong Guo, Feng Cheng, Zhixu Du +21
The rapid development of large language models (LLMs) has significantly transformed the field of artificial intelligence, demonstrating remarkable capabilities in natural language…
Exploiting Spin-Orbit Torque Devices as Reconfigurable Logic for Circuit Obfuscation
Jianlei Yang, Xueyan Wang, Qiang Zhou +5
Circuit obfuscation is a frequently used approach to conceal logic functionalities in order to prevent reverse engineering attacks on fabricated chips. Efficient obfuscation implem…
Spintronics based Stochastic Computing for Efficient Bayesian Inference System
Xiaotao Jia, Jianlei Yang, Zhaohao Wang +4
Bayesian inference is an effective approach for solving statistical learning problems especially with uncertainty and incompleteness. However, inference efficiencies are physically…
Thread Batching for High-performance Energy-efficient GPU Memory Design
Bing Li, Mengjie Mao, Xiaoxiao Liu +6
Massive multi-threading in GPU imposes tremendous pressure on memory subsystems. Due to rapid growth in thread-level parallelism of GPU and slowly improved peak memory bandwidth, t…
RED: A ReRAM-based Deconvolution Accelerator
Zichen Fan, Ziru Li, Bing Li +3
Deconvolution has been widespread in neural networks. For example, it is essential for performing unsupervised learning in generative adversarial networks or constructing fully con…
An Overview of In-memory Processing with Emerging Non-volatile Memory for Data-intensive Applications
Bing Li, Bonan Yan, Hai +1
The conventional von Neumann architecture has been revealed as a major performance and energy bottleneck for rising data-intensive applications. %, due to the intensive data moveme…
AutoRAC: Automated Processing-in-Memory Accelerator Design for Recommender Systems
Feng Cheng, Tunhou Zhang, Junyao Zhang +6
The performance bottleneck of deep-learning-based recommender systems resides in their backbone Deep Neural Networks. By integrating Processing-In-Memory~(PIM) architectures, resea…