5 papers · 1 filter
CATPO: Critique-Augmented Tree Policy Optimization
Ayush Singh, Umang Goyal, Ankur Dahiya
Reinforcement learning with verifiable rewards (RLVR) has become a dominant paradigm for improving the reasoning capabilities of large language models (LLMs). Recent tree-based met…
Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models
Pratham Singla, Shivank Garg, Ayush Singh +2
Recent advances in post-training techniques have endowed Large Language Models (LLMs) with enhanced capabilities for tackling complex, logic-intensive tasks through the generation…
IPO: Your Language Model is Secretly a Preference Classifier
Shivank Garg, Ayush Singh, Shweta Singh +1
Reinforcement learning from human feedback (RLHF) has emerged as the primary method for aligning large language models (LLMs) with human preferences. While it enables LLMs to achie…
LoRA-Mini : Adaptation Matrices Decomposition and Selective Training
Ayush Singh, Rajdeep Aher, Shivank Garg
The rapid advancements in large language models (LLMs) have revolutionized natural language processing, creating an increased need for efficient, task-specific fine-tuning methods.…
Are VLMs Really Blind
Ayush Singh, Mansi Gupta, Shivank Garg
Vision Language Models excel in handling a wide range of complex tasks, including Optical Character Recognition (OCR), Visual Question Answering (VQA), and advanced geometric reaso…