2 papers
cs.LG2026
ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization
Nirmal Patel, Fei Wang, Inderjit S. Dhillon
The alignment of Large Language Models (LLMs) utilizes Reinforcement Learning from AI Feedback (RLAIF) for non-verifiable domains such as long-form question answering and open-ende…
cs.AI2025
A Survey of Sustainability in Large Language Models: Applications, Economics, and Challenges
Aditi Singh, Nirmal Prakashbhai Patel, Abul Ehtesham +2
Large Language Models (LLMs) have transformed numerous domains by providing advanced capabilities in natural language understanding, generation, and reasoning. Despite their ground…