2 papers
cs.CL2026
PRISM: A Unified Framework for Post-Training LLMs Without Verifiable Rewards
Mukesh Ghimire, Aosong Feng, Liwen You +3
Current techniques for post-training Large Language Models (LLMs) rely either on costly human supervision or on external verifiers to boost performance on tasks such as mathematica…
cs.CR2025
TaeBench: Improving Quality of Toxic Adversarial Examples
Xuan Zhu, Dmitriy Bespalov, Liwen You +2
Toxicity text detectors can be vulnerable to adversarial examples - small perturbations to input text that fool the systems into wrong detection. Existing attack algorithms are tim…