From the 2 of 7 linked papers with an AI index.
4 papers · 1 filter
Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation
Genglin Liu, Muye Zhang, Krishnamurthy Viswanathan +5
The paper introduces an automated, multi‑agent framework that creates hard adversarial examples for multimodal large language models to improve content safety, achieving a signific…
PM-Bench: Evaluating Prospective Memory in LLM Agents
Genglin Liu, Saadia Gabriel
The paper introduces PM-Bench, a text-based benchmark that evaluates how well large language model agents can remember and act on future intentions while handling ongoing tasks.
WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance
Genglin Liu, Shijie Geng, Sha Li +4
Multimodal LLM-powered agents have recently demonstrated impressive capabilities in web navigation, enabling agents to complete complex browsing tasks across diverse domains. Howev…
SciCode: A Research Coding Benchmark Curated by Scientists
Minyang Tian, Luyu Gao, Shizhuo Dylan Zhang +27
Since language models (LMs) now outperform average humans on many challenging tasks, it has become increasingly difficult to develop challenging, high-quality, and realistic evalua…