From the 1 of 5 linked papers with an AI index.
1 paper · 1 filter
Genglin Liu, Saadia Gabriel
The paper introduces PM-Bench, a text-based benchmark that evaluates how well large language model agents can remember and act on future intentions while handling ongoing tasks.