6 papers
Mitigating Safety Tax via Distribution-Grounded Refinement in Large Reasoning Models
Yingsha Xie, Tiansheng Huang, Enneng Yang +5
Safety alignment incurs safety tax that perturbs a large reasoning model's (LRM) general reasoning ability. Existing datasets used for safety alignment for an LRM are usually const…
Empowering Reliable Visual-Centric Instruction Following in MLLMs
Weilei He, Feng Ju, Zhiyuan Fan +3
Evaluating the instruction-following (IF) capabilities of Multimodal Large Language Models (MLLMs) is essential for rigorously assessing how faithfully model outputs adhere to user…
Reasoning Path Divergence: A New Metric and Curation Strategy to Unlock LLM Diverse Thinking
Feng Ju, Zeyu Qin, Rui Min +3
While Test-Time Scaling (TTS) has proven effective in improving the reasoning ability of large language models (LLMs), low diversity in model outputs often becomes a bottleneck; th…
RemoteReasoner: Towards Unifying Geospatial Reasoning Workflow
Liang Yao, Fan Liu, Hongbo Lu +5
Remote sensing imagery presents vast, inherently unstructured spatial data, necessitating sophisticated reasoning to interpret complex user intents and contextual relationships bey…
Catch Me If You Can: How Smaller Reasoning Models Pretend to Reason with Mathematical Fidelity
Subramanyam Sahoo, Vinija Jain, Saanidhya Vats +4
Current evaluation of mathematical reasoning in language models relies primarily on answer accuracy, potentially masking fundamental failures in logical computation. We introduce a…
PADBen: A Comprehensive Benchmark for Evaluating AI Text Detectors Against Paraphrase Attacks
Yiwei Zha, Rui Min, Shanu Sushmita
While AI-generated text (AIGT) detectors achieve over 90\% accuracy on direct LLM outputs, they fail catastrophically against iteratively-paraphrased content. We investigate why it…