Showing 2026 · cs.AIShow all
2 papers · 2 filters
cs.AI2026
Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning
Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer +4
Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing returns: new rollouts often repeat existing ans…
cs.AI2026
Pressure, What Pressure? Sycophancy Disentanglement in Language Models via Reward Decomposition
Muhammad Ahmed Mohsin, Ahsan Bilal, Muhammad Umer +1
Large language models exhibit sycophancy, the tendency to shift their stated positions toward perceived user preferences or authority cues regardless of evidence. Standard alignmen…