2 papers
cs.CL2025
IPO: Your Language Model is Secretly a Preference Classifier
Shivank Garg, Ayush Singh, Shweta Singh +1
Reinforcement learning from human feedback (RLHF) has emerged as the primary method for aligning large language models (LLMs) with human preferences. While it enables LLMs to achie…
cs.MA2024
Adaptive Urban Planning: A Hybrid Framework for Balanced City Development
Pratham Singla, Ayush Singh, Adesh Gupta +1
Urban planning faces a critical challenge in balancing city-wide infrastructure needs with localized demographic preferences, particularly in rapidly developing regions. Although e…