Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets
Jack Hopkins, Dipika Khullar, Fabien Roger
Black box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information. To better elicit hidden information…
cs.AI2026
Self-Attribution Bias: When AI Monitors Go Easy on Themselves
Dipika Khullar, Jack Hopkins, Rowan Wang +1
Agentic systems increasingly rely on language models to monitor their own behavior. For example, coding agents may self critique generated code for pull request approval or assess…
cs.AI2026
Co-Evolving Agents: Learning from Failures as Hard Negatives
Yeonsung Jung, Trilok Padhi, Sina Shaham +4
The rapid progress of large foundation models has accelerated the development of task-specialized agents across diverse domains. However, the effectiveness of agents remains tightl…