16 papers
Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed
Haokun Lin, Kaijie Zhu, Haobo Xu +4
Small Language Models (SLMs) have emerged as a more efficient alternative to traditional Large Language Models (LLMs), offering promising potential in resource-constrained scenario…
Don't Regenerate, Debug: A Domain-Specific Agent for Repairing Near-Miss Hardware Operators
Yansong Sun, Shenxiu Wu, Siyuan Chen +6
Kernel generation for hardware accelerators such as GPUs and NPUs has become a proving ground for large language models (LLMs), and state-of-the-art systems raise correctness throu…
Fine-tuning Large Language Model for Automated Algorithm Design
Fei Liu, Rui Zhang, Xi Lin +2
The integration of large language models (LLMs) into automated algorithm design has shown promising potential. A prevalent approach embeds LLMs within search routines to iterativel…
DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization
Haokun Lin, Xinle Jia, Haobo Xu +7
The MXFP4 microscaling format, which partitions tensors into blocks of 32 elements sharing an E8M0 scaling factor, has emerged as a promising substrate for efficient LLM inference,…
Survey on Neural Routing Solvers
Yunpeng Ba, Xi Lin, Changliang Zhou +7
Neural routing solvers (NRSs) that leverage deep learning to tackle vehicle routing problems have demonstrated notable potential for practical applications. By learning implicit he…
Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
Haokun Lin, Haobo Xu, Yichen Wu +6
Recent advances in diffusion large language models (dLLMs) have introduced a promising alternative to autoregressive (AR) LLMs for natural language generation tasks, leveraging ful…