2 papers
cs.LG2025
KL-Regularized Reinforcement Learning is Designed to Mode Collapse
Anthony GX-Chen, Jatin Prakash, Jeff Guo +2
It is commonly believed that optimizing the reverse KL divergence results in "mode seeking", while optimizing forward KL results in "mass covering", with the latter being preferred…
cs.AI2025
Language Agents Mirror Human Causal Reasoning Biases. How Can We Help Them Think Like Scientists?
Anthony GX-Chen, Dongyan Lin, Mandana Samiei +4
Language model (LM) agents are increasingly used as autonomous decision-makers which need to actively gather information to guide their decisions. A crucial cognitive skill for suc…