9 papers
Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning
Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer +4
Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing returns: new rollouts often repeat existing ans…
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Xiaomin Li, Yuexing Hao, Jianheng Hou +90
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and inter…
Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation
Vishvesh Bhat, Jay Vaghasiya, Muhammad Ahmed Mohsin +1
Tool-calling benchmarks are increasingly used to rank language-model agents, yet their scores are often treated as ground truth without validating the evaluators themselves. We pre…
What If We Allocate Test-Time Compute Adaptively?
Ahsan Bilal, Ahmed Mohsin, Muhammad Umer +4
Test-time compute scaling allocates inference computation uniformly, uses fixed sampling strategies, and applies verification only for reranking. In contrast, we propose a verifier…
MAVEN: Improving Generalization in Agentic Tool Calling
Omkar Ghugarkar, Vishvesh Bhat, Muhammad Ahmed Mohsin +1
Generalization across agentic tool-calling environments remains a central challenge for reliable agentic reasoning systems. Although large language models achieve strong results on…
General Preference Reinforcement Learning
Muhammad Umer, Muhammad Ahmed Mohsin, Ahsan Bilal +5
Post-training has split large language model (LLM) alignment into two largely disconnected tracks. Online reinforcement learning (RL) with verifiable rewards drives emergent reason…