4 papers · 1 filter
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models
Yu Feng, Chunting Zang, Chen Shen +4
Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pairs. We audit two widely used safety-guard…
Purified OPSD: On-Policy Self-Distillation Without Losing How to Think
Zhanming Shen, Jintao Tong, Shaotian Yan +9
On-policy self-distillation (OPSD) has emerged as a promising paradigm for improving LLM reasoning, where a privileged teacher with access to reference solutions provides token-lev…
MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing
Bangbang Zhou, Hangdi Xing, Yifan Chen +8
Document parsing converts visually rich documents into machine-readable structured representations, forming a crucial foundation for information systems. Although many benchmarks h…
SalaMAnder: Shapley-based Mathematical Expression Attribution and Metric for Chain-of-Thought Reasoning
Yue Xin, Chen Shen, Shaotian Yan +5
Chain-of-Thought (CoT) prompting enhances the math reasoning capability of large language models (LLMs) to a large margin. However, the mechanism underlying such improvements remai…