From the 1 of 5 linked papers with an AI index.
1 citations · 1 across the 4 of their papers we have counts for
3 papers · 1 filter
Monadic Context Engineering
Yifan Zhang, Yang Yuan, Mengdi Wang +1
The proliferation of Large Language Models (LLMs) has catalyzed a shift towards autonomous agents capable of complex reasoning and tool use. However, current agent architectures ar…
Interactive Benchmarks
Baoqing Yue, Zihan Zhu, Yutong Han +6
Existing reasoning evaluation paradigms suffer from different limitations: fixed benchmarks are increasingly saturated and vulnerable to contamination, while preference-based evalu…
Web World Models
Jichen Feng, Yifan Zhang, Chenggong Zhang +3
Language agents increasingly require persistent worlds in which they can act, remember, and learn. Existing approaches sit at two extremes: conventional web frameworks provide reli…