2 papers
cs.AI2026
PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models
Zetian Ouyang, Linlin Wang, Gerard de Melo +1
Despite the pivotal role of numerical reasoning as the cornerstone of mathematical capabilities in large language models (LLMs) across applications, few benchmarks evaluate LLMs by…
cs.CR2025
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks
Xin Yi, Yue Li, Dongsheng Shi +3
Ensuring safety alignment is a critical requirement for large language models (LLMs), particularly given increasing deployment in real-world applications. Despite considerable adva…