Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Reasoning Path Divergence: A New Metric and Curation Strategy to Unlock LLM Diverse Thinking
Feng Ju, Zeyu Qin, Rui Min +3
While Test-Time Scaling (TTS) has proven effective in improving the reasoning ability of large language models (LLMs), low diversity in model outputs often becomes a bottleneck; th…
cs.CL2025
Catch Me If You Can: How Smaller Reasoning Models Pretend to Reason with Mathematical Fidelity
Subramanyam Sahoo, Vinija Jain, Saanidhya Vats +4
Current evaluation of mathematical reasoning in language models relies primarily on answer accuracy, potentially masking fundamental failures in logical computation. We introduce a…
cs.CL2025
PADBen: A Comprehensive Benchmark for Evaluating AI Text Detectors Against Paraphrase Attacks
Yiwei Zha, Rui Min, Shanu Sushmita
While AI-generated text (AIGT) detectors achieve over 90\% accuracy on direct LLM outputs, they fail catastrophically against iteratively-paraphrased content. We investigate why it…