8 papers
PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models
Zetian Ouyang, Linlin Wang, Gerard de Melo +1
Despite the pivotal role of numerical reasoning as the cornerstone of mathematical capabilities in large language models (LLMs) across applications, few benchmarks evaluate LLMs by…
Disentangling Language Roles in Multilingual LLM Task Execution
Qishi Zhan, Minxuan Hu, Seoyeon Jang +7
Multilingual LLMs are increasingly used when instruction, source content, and required response languages do not coincide. Existing benchmarks have expanded multilingual instructio…
Unified Defense for Large Language Models against Jailbreak and Fine-Tuning Attacks in Education
Xin Yi, Yue Li, Dongsheng Shi +3
Large Language Models (LLMs) are increasingly integrated into educational applications. However, they remain vulnerable to jailbreak and fine-tuning attacks, which can compromise s…
Unified attacks to large language model watermarks: spoofing and scrubbing in unauthorized knowledge distillation
Xin Yi, Yue Li, Shunfan Zheng +3
Watermarking has emerged as a critical technique for combating misinformation and protecting intellectual property in large language models (LLMs). A recent discovery, termed water…
Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks
Xin Yi, Yue Li, Dongsheng Shi +3
Ensuring safety alignment is a critical requirement for large language models (LLMs), particularly given increasing deployment in real-world applications. Despite considerable adva…
AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation
Xiechi Zhang, Zetian Ouyang, Linlin Wang +6
With the proliferation of large language models (LLMs) in the medical domain, there is increasing demand for improved evaluation techniques to assess their capabilities. However, t…