3 citations · 3 across the 7 of their papers we have counts for
4 papers · 1 filter
Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation
Genglin Liu, Muye Zhang, Krishnamurthy Viswanathan +5
Multimodal Large Language Models (MLLMs) are increasingly deployed for nuanced content safety and moderation tasks, yet they remain vulnerable to adversarial attacks and out-of-dis…
PM-Bench: Evaluating Prospective Memory in LLM Agents
Genglin Liu, Saadia Gabriel
A significant challenge in agentic AI is prospective memory: the ability to execute an intention at a specific future cue or state while other activities are ongoing. We introduce…
WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance
Genglin Liu, Shijie Geng, Sha Li +4
Multimodal LLM-powered agents have recently demonstrated impressive capabilities in web navigation, enabling agents to complete complex browsing tasks across diverse domains. Howev…
SciCode: A Research Coding Benchmark Curated by Scientists
Minyang Tian, Luyu Gao, Shizhuo Dylan Zhang +27
Since language models (LMs) now outperform average humans on many challenging tasks, it has become increasingly difficult to develop challenging, high-quality, and realistic evalua…