3 papers
cs.LG2026
StableQAT: Stable Quantization-Aware Training at Ultra-Low Bitwidths
Tianyi Chen, Sihan Chen, Xiaoyi Qu +5
Quantization-aware training (QAT) is essential for deploying large models under strict memory and latency constraints, yet achieving stable and robust optimization at ultra-low bit…
cs.LG2026
WINA: Weight Informed Neuron Activation for Accelerating Large Language Model Inference
Sihan Chen, Dan Zhao, Jongwoo Ko +5
The growing computational demands of large language models (LLMs) make efficient inference and activation strategies increasingly critical. While recent approaches, such as Mixture…
cs.CV2026
LightFormer: A lightweight and efficient decoder for remote sensing image segmentation
Sihang Chen, Lijun Yun, Ze Liu +4
Deep learning techniques have achieved remarkable success in the semantic segmentation of remote sensing images and in land-use change detection. Nevertheless, their real-time depl…