collaborators

7 papers

cs.AI2026

CONTRAMEM: Learning Self-Evolving Procedural Memory from Contrasting Multi-Model Trajectories

Zheyuan Deng, Binghang Lu, Hanqi Feng +10

Autonomous computer-use agents are increasingly applied to long-horizon tasks requiring coordinated application calls, persistent state tracking, and verifier-sensitive writes, yet…

cs.CL2026

MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations

Sky Ng, Brihi Joshi, Ishan Gupta +47

Long-horizon, multi-agent language model (LM) simulations are widely proposed for studying social behavior, yet instruments to measure whether persona-conditioned agents maintain i…

cs.HC2026

PersonaEval: Persona-Based User Simulation for Evaluating Interactive Applications

Yifan Simon Liu, Qianfeng Wen, Yilan Fan +40

Real user studies are important for understanding how people interact with systems under test or already deployed. In practice, however, they are often costly, time-consuming, and…

cs.AI2026

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

Xiaomin Li, Yuexing Hao, Jianheng Hou +90

Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and inter…

cs.AI2026

BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance

Harmon Bhasin, Kevin Flyangolts, Dianzhuo Wang +9

As pathogen genomic surveillance scales, the bottleneck is shifting from data generation to analysis. We present BioSecBench-Surveillance, a verifiable benchmark of 100 evaluations…

cs.CR2026

BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment

Edwin H. Wintermute, Harmon Bhasin, Christina M. Agapakis +13

As AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse. We present BioSecBench-Refusal, a benchmark for risk iden…