4 citations · 4 across the 2 of their papers we have counts for
4 papers · 1 filter
MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
Zexue He, Yu Wang, Churan Zhi +11
Existing evaluations of agents with memory typically assess memorization and action in isolation. One class of benchmarks evaluates memorization by testing recall of past conversat…
ChemAgent: Self-updating Library in Large Language Models Improves Chemical Reasoning
Xiangru Tang, Tianyu Hu, Muyang Ye +9
Chemical reasoning usually involves complex, multi-step processes that demand precise calculations, where even minor errors can lead to cascading failures. Furthermore, large langu…
Smoothing Dialogue States for Open Conversational Machine Reading
Zhuosheng Zhang, Siru Ouyang, Hai Zhao +2
Conversational machine reading (CMR) requires machines to communicate with humans through multi-turn interactions between two salient dialogue states of decision making and questio…
Dialogue Graph Modeling for Conversational Machine Reading
Siru Ouyang, Zhuosheng Zhang, Hai Zhao
Conversational Machine Reading (CMR) aims at answering questions in a complicated manner. Machine needs to answer questions through interactions with users based on given rule docu…