6 papers
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
Jinming Lu, Jiayi Tian, Hai Li +2
The increasing demand for on-device training of deep neural networks (DNNs) aims to leverage personal data for high-performance applications while addressing privacy concerns and r…
FLAT-LLM: Fine-grained Low-rank Activation Space Transformation for Large Language Model Compression
Jiayi Tian, Ryan Solgi, Jinming Lu +3
Large Language Models (LLMs) have enabled remarkable progress in natural language processing, yet their high computational and memory demands pose challenges for deployment in reso…
Tensor-Compressed and Fully-Quantized Training of Neural PDE Solvers
Jinming Lu, Jiayi Tian, Yequan Zhao +2
Physics-Informed Neural Networks (PINNs) have emerged as a promising paradigm for solving partial differential equations (PDEs) by embedding physical laws into neural network train…
DeepOHeat-v1: Efficient Operator Learning for Fast and Trustworthy Thermal Simulation and Optimization in 3D-IC Design
Xinling Yu, Ziyue Liu, Hai Li +5
Thermal analysis is crucial in 3D-IC design due to increased power density and complex heat dissipation paths. Although operator learning frameworks such as DeepOHeat~\cite{liu2023…
Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
Jiayi Tian, Jinming Lu, Hai Li +4
Transformer models have achieved state-of-the-art performance across a wide range of machine learning tasks. There is growing interest in training transformers on resource-constrai…
Poor Man's Training on MCUs: A Memory-Efficient Quantized Back-Propagation-Free Approach
Yequan Zhao, Hai Li, Ian Young +1
Back propagation (BP) is the default solution for gradient computation in neural network training. However, implementing BP-based training on various edge devices such as FPGA, mic…