4 papers
Max-Window Scale Estimation for Near-Lossless HiF8 W8A8 Quantization-Aware Training
Yingying Cheng, Jinquan Shi, Li Zhou +4
Quantization-aware training (QAT) with low-bit floating-point formats enables efficient LLM deployment, yet introduces subtle failure modes invisible to standard training metrics.…
SDSL-Solver: Scalable Distributed Sparse Linear Solvers for Large-Scale Interior Point Methods
Shaofeng Yang, Yunting Wang, Yingying Cheng +3
The solution of sparse linear systems constitutes the dominant computational bottleneck in interior point methods (IPMs), frequently consuming over 70% of the total solution time.…
Making Slow Thinking Faster: Compressing LLM Chain-of-Thought via Step Entropy
Zeju Li, Jianyuan Zhong, Ziyang Zheng +5
Large Language Models (LLMs) using Chain-of-Thought (CoT) prompting excel at complex reasoning but generate verbose thought processes with considerable redundancy, leading to incre…
GridCodex: A RAG-Driven AI Framework for Power Grid Code Reasoning and Compliance
Jinquan Shi, Yingying Cheng, Fan Zhang +3
The global shift towards renewable energy presents unprecedented challenges for the electricity industry, making regulatory reasoning and compliance increasingly vital. Grid codes,…