3 papers
cs.AI2026
EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models
Xinning Li, Kemunto Ochwang'i, Aryasomayajula Ram Bharadwaj +2
Frontier large language models can often recognize when they are being evaluated, a capability known as evaluation awareness. If models behave differently in evaluations than in de…
cs.CL2025
Adaptive Activation Steering for Efficient LLM Reasoning via Closed-Loop PID Control
Aryasomayajula Ram Bharadwaj
Reasoning LLMs trained with long chain-of-thought often overthink: they spend tokens on redundant reflection and transitions that inflate cost without improving accuracy. Static ac…
cs.CL2024
Understanding Hidden Computations in Chain-of-Thought Reasoning
Aryasomayajula Ram Bharadwaj
Chain-of-Thought (CoT) prompting has significantly enhanced the reasoning abilities of large language models. However, recent studies have shown that models can still perform compl…