Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment
Qiuyu Tian, Haojie Yin, Yingce Xia +2
AI research often requires decisions before future evidence exists: which bottleneck to attack, which direction to pursue, or where a project should be positioned. We introduce For…
cs.AI2026
OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories
Yibing Liu, Yangze Liu, Xiaolong Yin +4
Task success can hide process anomalies in real-world agent executions. An agent may pass the final task oracle while still accumulating unresolved ambiguity, unsafe external write…