2 papers
cs.LG2026
VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation
Yifan Bai, Xiaoyang Liu, Zihao Mou +7
As large language models (LLMs) are increasingly deployed for software engineering, constructing high-quality benchmarks is crucial for evaluating not just the functional correctne…
cs.CL2026
Select to Think: Unlocking SLM Potential with Local Sufficiency
Wenxuan Ye, Yangyang Zhang, Xueli An +2
Small language models (SLMs) offer efficient deployment, yet they often lag behind their larger counterparts (LLMs) in reasoning. Existing remedies either invoke an LLM at points o…