4 papers
When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models
Sailesh Panda, Pritam Kadasi, Abhishek Upperwal +1
Large language models (LLMs) often achieve strong performance on reasoning benchmarks, but final-answer accuracy alone does not show whether they faithfully execute the procedure s…
Eka-Eval: An Evaluation Framework for Low-Resource Multilingual Large Language Models
Samridhi Raj Sinha, Rajvee Sheth, Abhishek Upperwal +1
The rapid evolution of Large Language Models' has underscored the need for evaluation frameworks that are globally applicable, flexible, and modular, and that support a wide range…
Task--Specificity Score: Measuring How Much Instructions Really Matter for Supervision
Pritam Kadasi, Abhishek Upperwal, Mayank Singh
Instruction tuning is now the default way to train and adapt large language models, but many instruction--input--output pairs are only weakly specified: for a given input, the same…
ADAPT: Learning Task Mixtures for Budget-Constrained Instruction Tuning
Pritam Kadasi, Abhishek Upperwal, Mayank SIngh
We propose ADAPT, a meta-learning algorithm that \emph{learns} task sampling proportions under an explicit token budget for multi-task instruction tuning. Instead of fixing task we…