16 papers
Benchmarking the Benchmarks: A Validity Audit of Tool-Calling Evaluation
Vishvesh Bhat, Jay Vaghasiya, Muhammad Ahmed Mohsin +1
Tool-calling benchmarks are increasingly used to rank language-model agents, yet their scores are often treated as ground truth without validating the evaluators themselves. We pre…
MAVEN: Improving Generalization in Agentic Tool Calling
Omkar Ghugarkar, Vishvesh Bhat, Muhammad Ahmed Mohsin +1
Generalization across agentic tool-calling environments remains a central challenge for reliable agentic reasoning systems. Although large language models achieve strong results on…
Training-Free Adaptation of New-Generation LLMs using Legacy Clinical Models
Sasha Ronaghi, Chloe Stanwyck, Asad Aali +4
Adapting language models to the clinical domain through continued pretraining and instruction tuning requires costly retraining for each new model generation. We propose Cross-Arch…
: Stratified Scaling Search for Test-Time in Diffusion Language Models
Ahsan Bilal, Muhammad Ahmed Mohsin, Muhammad Umer +6
Test-time scaling investigates whether a fixed diffusion language model (DLM) can generate better outputs when given more inference compute, without additional training. However, n…
Structured Prompts Improve Evaluation of Language Models
Asad Aali, Muhammad Ahmed Mohsin, Vasiliki Bikia +15
As language models (LMs) are increasingly adopted across domains, high-quality benchmarking frameworks are essential for guiding deployment decisions. In practice, however, framewo…
Clinical Note Bloat Reduction for Efficient LLM Use
Jordan L. Cahoon, Chloe Stanwyck, Asad Aali +5
Health systems are rapidly deploying large language models (LLMs) that use clinical notes for clinical decision support applications. However, modern documentation practices rely h…