3 papers
cs.IR2026
LoopCTR: Unlocking the Loop Scaling Power for Click-Through Rate Prediction
Jiakai Tang, Runfeng Zhang, Weiqiu Wang +7
Scaling Transformer-based click-through rate (CTR) models by stacking more parameters brings growing computational and storage overhead, creating a widening gap between scaling amb…
cs.LG2026
LBLLM: Lightweight Binarization of Large Language Models via Three-Stage Distillation
Siqing Song, Chuang Wang, Yong Lang +2
Deploying large language models (LLMs) in resource-constrained environments is hindered by heavy computational and memory requirements. We present LBLLM, a lightweight binarization…
cs.LG2025
Achieving binary weight and activation for LLMs using Post-Training Quantization
Siqing Song, Chuang Wang, Ruiqi Wang +2
Quantizing large language models (LLMs) to 1-bit precision significantly reduces computational costs, but existing quantization techniques suffer from noticeable performance degrad…