Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Test-Time Scaling for Scientific Equation Discovery
Haowei Lin, Hubert Lim, Xiangyu Wang +2
Test-time scaling (TTS) improves language model reasoning by allocating additional test-time compute, but prior work mainly studies closed-ended tasks such as math and coding. We s…
cs.CL2024
WalledEval: A Comprehensive Safety Evaluation Toolkit for Large Language Models
Prannaya Gupta, Le Qi Yau, Hao Han Low +8
WalledEval is a comprehensive AI safety testing toolkit designed to evaluate large language models (LLMs). It accommodates a diverse range of models, including both open-weight and…