actor-critic PPO 1confidence estimation 1large language models 1quantitative prediction 1reinforcement learning 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CL2026
From Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence Estimation
Mehak Dhaliwal, Rasta Tadayon, Andong Hua +2
The paper introduces CARE-PPO, a reinforcement‑learning framework that fine‑tunes large language models to make accurate numeric predictions while simultaneously learning confidenc…
cs.CL2026
English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training
Mehak Dhaliwal, Shashwat Chaurasia, Yao Qin +2
Despite the widespread multilingual deployment of large language models, post-training pipelines remain predominantly English-centric, contributing to performance disparities acros…
cs.CL2026
NutriBench: A Dataset for Evaluating Large Language Models on Nutrition Estimation from Meal Descriptions
Andong Hua, Mehak Preet Dhaliwal, Laya Pullela +2
Accurate nutrition estimation helps people make informed dietary choices and is essential in the prevention of serious health complications. We present NutriBench, the first public…