Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
From 0-Order Selection to 2-Order Judgment: Combinatorial Hardening Exposes Compositional Failures in Frontier LLMs
Hanmeng Liu, Shichao Weng, Xiulai Liu +3
Multiple-choice reasoning benchmarks face dual challenges: rapid saturation from advancing models and data contamination that undermines static evaluations. Ad-hoc hardening method…
cs.CL2025
GLoRE: Evaluating Logical Reasoning of Large Language Models
Hanmeng liu, Zhiyang Teng, Ruoxi Ning +4
Large language models (LLMs) have shown significant general language understanding abilities. However, there has been a scarcity of attempts to assess the logical reasoning capacit…
cs.CL2024
Break the Chain: Large Language Models Can be Shortcut Reasoners
Mengru Ding, Hanmeng Liu, Zhizhang Fu +3
Recent advancements in Chain-of-Thought (CoT) reasoning utilize complex modules but are hampered by high token consumption, limited applicability, and challenges in reproducibility…