collaborators

5 papers

cs.LG2026

Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks

Lawrence Keunho Jang, Jing Yu Koh, Daniel Fried +1

Existing web agent benchmarks have largely converged on short, single-site tasks that frontier models are approaching saturation on. However, real world web use consists of long-ho…

cs.AI2026

Agent psychometrics: Task-level performance prediction in agentic coding benchmarks

Chris Ge, Daria Kryvosheieva, Daniel Fried +2

As the focus in LLM-based coding shifts from static single-step code generation to multi-step agentic interaction with tools and environments, understanding which tasks will challe…

cs.CL2026

Generative Value Conflicts Reveal LLM Priorities

Andy Liu, Kshitish Ghate, Mona Diab +3

Past work seeks to align large language model (LLM)-based assistants with a target set of values, but such assistants are frequently forced to make tradeoffs between values when de…

cs.CY2025

Measuring Fine-Grained Negotiation Tactics of Humans and LLMs in Diplomacy

Wenkai Li, Lynnette Hui Xian Ng, Andy Liu +1

The study of negotiation styles dates back to Aristotle's ethos-pathos-logos rhetoric. Prior efforts primarily studied the success of negotiation agents. Here, we shift the focus t…

cs.MA2025

Dynamic Coalition Structure Detection in Natural Language-based Interactions

Abhishek N. Kulkarni, Andy Liu, Jean-Raphael Gaglione +2

In strategic multi-agent sequential interactions, detecting dynamic coalition structures is crucial for understanding how self-interested agents coordinate to influence outcomes. H…