6 papers
Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation
Genglin Liu, Muye Zhang, Krishnamurthy Viswanathan +5
The paper introduces an automated, multi‑agent framework that creates hard adversarial examples for multimodal large language models to improve content safety, achieving a signific…
PM-Bench: Evaluating Prospective Memory in LLM Agents
Genglin Liu, Saadia Gabriel
The paper introduces PM-Bench, a text-based benchmark that evaluates how well large language model agents can remember and act on future intentions while handling ongoing tasks.
WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance
Genglin Liu, Shijie Geng, Sha Li +4
Multimodal LLM-powered agents have recently demonstrated impressive capabilities in web navigation, enabling agents to complete complex browsing tasks across diverse domains. Howev…
AI Debate Aids Assessment of Controversial Claims
Salman Rahman, Sheriff Issaka, Ashima Suvarna +11
As AI grows more powerful, it will increasingly shape how we understand the world. But with this influence comes the risk of amplifying misinformation and deepening social divides-…
MOSAIC: Modeling Social AI for Content Dissemination and Regulation in Multi-Agent Simulations
Genglin Liu, Vivian Le, Salman Rahman +3
We present a novel, open-source social network simulation framework, MOSAIC, where generative language agents predict user behaviors such as liking, sharing, and flagging content.…
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
Salman Rahman, Liwei Jiang, James Shiffer +7
Multi-turn interactions with language models (LMs) pose critical safety risks, as harmful intent can be strategically spread across exchanges. Yet, the vast majority of prior work…