Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning
Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer +4
Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing returns: new rollouts often repeat existing ans…
cs.AI2026
Pressure, What Pressure? Sycophancy Disentanglement in Language Models via Reward Decomposition
Muhammad Ahmed Mohsin, Ahsan Bilal, Muhammad Umer +1
Large language models exhibit sycophancy, the tendency to shift their stated positions toward perceived user preferences or authority cues regardless of evidence. Standard alignmen…
cs.AI2025
Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey
Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer +2
This survey explores the development of meta-thinking capabilities in Large Language Models (LLMs) from a Multi-Agent Reinforcement Learning (MARL) perspective. Meta-thinking self-…