1 paper
Huan Zhang, Mingju Chen, Dongxu Zhou +5
Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult…