3 papers
cs.CL2025
The Silent Judge: Unacknowledged Shortcut Bias in LLM-as-a-Judge
Arash Marioriyad, Mohammad Hossein Rohban, Mahdieh Soleymani Baghshah
Large language models (LLMs) are increasingly deployed as automatic judges to evaluate system outputs in tasks such as summarization, dialogue, and creative writing. A faithful jud…
cs.CL2025
Unspoken Hints: Accuracy Without Acknowledgement in LLM Reasoning
Arash Marioriyad, Shaygan Adim, Nima Alighardashi +2
Large language models (LLMs) increasingly rely on chain-of-thought (CoT) prompting to solve mathematical and logical reasoning tasks. Yet, a central question remains: to what exten…
cs.AI2025
Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization
Mohammad Mahdi Samiei Paqaleh, Arash Marioriyad, Arman Tahmasebi-Zadeh +3
Recent progress has pushed AI frontiers from pattern recognition tasks toward problems that require step by step, System2 style reasoning, especially with large language models. Ye…