Hardware for Machine Learning: Challenges and Opportunities
arXiv:1612.07625 · doi:10.1109/CICC.2017.7993626
Abstract
Machine learning plays a critical role in extracting meaningful information out of the zetabytes of sensor data collected every day. For some applications, the goal is to analyze and understand the data to identify trends (e.g., surveillance, portable/wearable electronics); in other applications, the goal is to take immediate action based the data (e.g., robotics/drones, self-driving cars, smart Internet of Things). For many of these applications, local embedded processing near the sensor is preferred over the cloud due to privacy or latency concerns, or limitations in the communication bandwidth. However, at the sensor there are often stringent constraints on energy consumption and cost in addition to throughput and accuracy requirements. Furthermore, flexibility is often required such that the processing can be adapted for different applications or environments (e.g., update the weights and model in the classifier). In many applications, machine learning often involves transforming the input data into a higher dimensional space, which, along with programmable weights, increases data movement and consequently energy consumption. In this paper, we will discuss how these challenges can be addressed at various levels of hardware design ranging from architecture, hardware-friendly algorithms, mixed-signal circuits, and advanced technologies (including memories and sensors).
Published as an invited conference paper at CICC 2017
References in corpus (1)
Cited by in corpus (11)
- FPGA-based Accelerators of Deep Learning Networks for Learning and Classification: A Review
- DORY: Automatic End-to-End Deployment of Real-World DNNs on Low-Cost IoT MCUs
- Efficient Hardware Acceleration of Sparsely Active Convolutional Spiking Neural Networks
- Binary classification of proteins by a Machine Learning approach
- CryoCiM: Cryogenic Compute-in-Memory based on the Quantum Anomalous Hall Effect
- Software Engineering Approaches for TinyML based IoT Embedded Vision: A Systematic Literature Review
- Single-Shot Matrix-Matrix Multiplication Optical Tensor Processor for Deep Learning
- LIBRA: Enabling Workload-aware Multi-dimensional Network Topology Optimization for Distributed Training of Large AI Models
- Data Streaming and Traffic Gathering in Mesh-based NoC for Deep Neural Network Acceleration
- How Do Companies Manage the Environmental Sustainability of AI? An Interview Study About Green AI Efforts and Regulations
- On the Impact of Partial Sums on Interconnect Bandwidth and Memory Accesses in a DNN Accelerator