actor-critic PPO 1confidence estimation 1large language models 1quantitative prediction 1reinforcement learning 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CL2026
From Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence Estimation
Mehak Dhaliwal, Rasta Tadayon, Andong Hua +2
The paper introduces CARE-PPO, a reinforcement‑learning framework that fine‑tunes large language models to make accurate numeric predictions while simultaneously learning confidenc…
cs.HC2025
Learning to Lie: Reinforcement Learning Attacks Damage Human-AI Teams and Teams of LLMs
Abed Kareem Musaffar, Anand Gokhale, Sirui Zeng +4
As artificial intelligence (AI) assistants become more widely adopted in safety-critical domains, it becomes important to develop safeguards against potential failures or adversari…