7 papers
CATPO: Critique-Augmented Tree Policy Optimization
Ayush Singh, Umang Goyal, Ankur Dahiya
Reinforcement learning with verifiable rewards (RLVR) has become a dominant paradigm for improving the reasoning capabilities of large language models (LLMs). Recent tree-based met…
Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models
Pratham Singla, Shivank Garg, Ayush Singh +2
Recent advances in post-training techniques have endowed Large Language Models (LLMs) with enhanced capabilities for tackling complex, logic-intensive tasks through the generation…
SIDiffAgent: Self-Improving Diffusion Agent
Shivank Garg, Ayush Singh, Gaurav Kumar Nayak
Text-to-image diffusion models have revolutionized generative AI, enabling high-quality and photorealistic image synthesis. However, their practical deployment remains hindered by…
IPO: Your Language Model is Secretly a Preference Classifier
Shivank Garg, Ayush Singh, Shweta Singh +1
Reinforcement learning from human feedback (RLHF) has emerged as the primary method for aligning large language models (LLMs) with human preferences. While it enables LLMs to achie…
Adaptive Urban Planning: A Hybrid Framework for Balanced City Development
Pratham Singla, Ayush Singh, Adesh Gupta +1
Urban planning faces a critical challenge in balancing city-wide infrastructure needs with localized demographic preferences, particularly in rapidly developing regions. Although e…
LoRA-Mini : Adaptation Matrices Decomposition and Selective Training
Ayush Singh, Rajdeep Aher, Shivank Garg
The rapid advancements in large language models (LLMs) have revolutionized natural language processing, creating an increased need for efficient, task-specific fine-tuning methods.…