6 papers
Tensor-Compressed and Fully-Quantized Training of Neural PDE Solvers
Jinming Lu, Jiayi Tian, Yequan Zhao +2
Physics-Informed Neural Networks (PINNs) have emerged as a promising paradigm for solving partial differential equations (PDEs) by embedding physical laws into neural network train…
SkipKV: Selective Skipping of KV Generation and Storage for Efficient Inference with Large Reasoning Models
Jiayi Tian, Seyedarmin Azizi, Yequan Zhao +7
Large reasoning models (LRMs) often incur significant key-value (KV) cache overhead, due to their linear growth with the verbose chain-of-thought (CoT) reasoning. This incurs both…
Comprehensive Design Space Exploration for Tensorized Neural Network Hardware Accelerators
Jinsong Zhang, Minghe Li, Jiayi Tian +2
High-order tensor decomposition has been widely adopted to obtain compact deep neural networks for edge deployment. However, existing studies focus primarily on its algorithmic adv…
FLAT-LLM: Fine-grained Low-rank Activation Space Transformation for Large Language Model Compression
Jiayi Tian, Ryan Solgi, Jinming Lu +3
Large Language Models (LLMs) have enabled remarkable progress in natural language processing, yet their high computational and memory demands pose challenges for deployment in reso…
FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training
Jinming Lu, Jiayi Tian, Hai Li +2
The increasing demand for on-device training of deep neural networks (DNNs) aims to leverage personal data for high-performance applications while addressing privacy concerns and r…
Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization
Jiayi Tian, Jinming Lu, Hai Li +4
Transformer models have achieved state-of-the-art performance across a wide range of machine learning tasks. There is growing interest in training transformers on resource-constrai…