4 papers
On the Practice of Scaling Search Conversion Rate Prediction
James Pak, Jyun-Yu Jiang, Fan Zhang +13
Scaling a Search Conversion Rate (CVR) prediction model, especially in high-traffic environments, presents a challenge: superior model quality needs to be balanced with strict cons…
Max-Window Scale Estimation for Near-Lossless HiF8 W8A8 Quantization-Aware Training
Yingying Cheng, Jinquan Shi, Li Zhou +4
Quantization-aware training (QAT) with low-bit floating-point formats enables efficient LLM deployment, yet introduces subtle failure modes invisible to standard training metrics.…
Making Slow Thinking Faster: Compressing LLM Chain-of-Thought via Step Entropy
Zeju Li, Jianyuan Zhong, Ziyang Zheng +5
Large Language Models (LLMs) using Chain-of-Thought (CoT) prompting excel at complex reasoning but generate verbose thought processes with considerable redundancy, leading to incre…
GridCodex: A RAG-Driven AI Framework for Power Grid Code Reasoning and Compliance
Jinquan Shi, Yingying Cheng, Fan Zhang +3
The global shift towards renewable energy presents unprecedented challenges for the electricity industry, making regulatory reasoning and compliance increasingly vital. Grid codes,…