papers

Publications (8)

cs.LG2026

FlashFPS: Efficient Farthest Point Sampling for Large-Scale Point Clouds via Pruning and Caching

Yuzhe Fu, Hancheng Ye, Cong Guo +7

Point-based Neural Networks (PNNs) have become a key approach for point cloud processing. However, a core operation in these models, Farthest Point Sampling (FPS), often introduces…

cs.AR2024

A Survey: Collaborative Hardware and Software Design in the Era of Large Language Models

Cong Guo, Feng Cheng, Zhixu Du +21

The rapid development of large language models (LLMs) has significantly transformed the field of artificial intelligence, demonstrating remarkable capabilities in natural language…

cs.ET2018

Exploiting Spin-Orbit Torque Devices as Reconfigurable Logic for Circuit Obfuscation

Jianlei Yang, Xueyan Wang, Qiang Zhou +5

Circuit obfuscation is a frequently used approach to conceal logic functionalities in order to prevent reverse engineering attacks on fabricated chips. Efficient obfuscation implem…

cs.ET2017

Spintronics based Stochastic Computing for Efficient Bayesian Inference System

Xiaotao Jia, Jianlei Yang, Zhaohao Wang +4

Bayesian inference is an effective approach for solving statistical learning problems especially with uncertainty and incompleteness. However, inference efficiencies are physically…

cs.AR2019

Thread Batching for High-performance Energy-efficient GPU Memory Design

Bing Li, Mengjie Mao, Xiaoxiao Liu +6

Massive multi-threading in GPU imposes tremendous pressure on memory subsystems. Due to rapid growth in thread-level parallelism of GPU and slowly improved peak memory bandwidth, t…

cs.ET2019

RED: A ReRAM-based Deconvolution Accelerator

Zichen Fan, Ziru Li, Bing Li +3

Deconvolution has been widespread in neural networks. For example, it is essential for performing unsupervised learning in generative adversarial networks or constructing fully con…

cs.AR2019

An Overview of In-memory Processing with Emerging Non-volatile Memory for Data-intensive Applications

Bing Li, Bonan Yan, Hai +1

The conventional von Neumann architecture has been revealed as a major performance and energy bottleneck for rising data-intensive applications. %, due to the intensive data moveme…

cs.AR2025

AutoRAC: Automated Processing-in-Memory Accelerator Design for Recommender Systems

Feng Cheng, Tunhou Zhang, Junyao Zhang +6

The performance bottleneck of deep-learning-based recommender systems resides in their backbone Deep Neural Networks. By integrating Processing-In-Memory~(PIM) architectures, resea…