2 papers
cs.LG2026
LLMForge: Multi-Backend Hardware-Aware Neural Architecture Search with Infinite-Head Attention for Edge Language Models
Xinting Jiang, Junyi Luo, Ruichen Qi +4
Sub-billion-parameter Transformer language models are increasingly deployed on edge devices, where the privacy, latency, and operating-cost advantages of on-device inference are co…
cs.AR2024
ConSmax: Hardware-Friendly Alternative Softmax with Learnable Parameters
Shiwei Liu, Guanchen Tao, Yifei Zou +7
The self-attention mechanism distinguishes transformer-based large language models (LLMs) apart from convolutional and recurrent neural networks. Despite the performance improvemen…