3 papers
cs.DC2026
FPTC: A Fast Parallel Transform-based Codec for Efficient Asymmetric Signal Compression
Ben Mechels, Ryan Billmeyer, Alexander Chen +2
Modern high-performance computing and Internet-of-Things deployments increasingly generate large volumes of signal data that must be compressed efficiently on resource-constrained…
cs.LG2026
GSR-GNN: Training Acceleration and Memory-Saving Framework of Deep GNNs on Circuit Graph
Yuebo Luo, Shiyang Li, Yifei Feng +3
Graph Neural Networks (GNNs) show strong promise for circuit analysis, but scaling to modern large-scale circuit graphs is limited by GPU memory and training cost, especially for d…
cs.LG2025
End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
Qitao Tan, Xiaoying Song, Jin Lu +9
Quantization is an effective technique to reduce the deployment cost of large language models (LLMs), and post-training quantization (PTQ) has been widely studied due to its effici…