Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Code Comprehension then Auditing for Unsupervised LLM Evaluation
Bhrij Patel, Souradip Chakraborty, Mengdi Wang +2
Large Language Models (LLMs) for unsupervised code correctness evaluation have recently gained attention because they can judge if code runs as intended without requiring reference…
cs.AI2025
Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models
Soumya Suvra Ghosal, Souradip Chakraborty, Avinash Reddy +6
Recent trends in test-time scaling for reasoning models (e.g., OpenAI o1, DeepSeek R1) have led to a popular belief that extending thinking traces using prompts like "Wait" or "Let…