1 paper · 1 filter
Huan Zhang, Mingju Chen, Dongxu Zhou +5
Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult…