From the 1 of 86 linked papers with an AI index.
3 citations · 6 across the 25 of their papers we have counts for
19 papers · 1 filter
BehaviorBench: Modeling Real-World User Decisions from Behavioral Traces
Liangwei Yang, Jielin Qiu, Zixiang Chen +9
Many decision-support settings require systems that adapt to individual users, but evaluation data for this problem remain limited. Existing benchmarks for user understanding often…
Counterparty Modeling is Not Strategy: The Limits of LLM Negotiators
Romain Cosentino, Sarath Shekkizhar, Adam Earle +1
Negotiation requires more than inferring what the other side wants: it requires using that information to make advantageous offers and counteroffers over multiple turns. We study w…
A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems
Zixuan Ke, Fangkai Jiao, Yifei Ming +9
Reasoning is a fundamental cognitive process that enables logical inference, problem-solving, and decision-making. With the rapid advancement of large language models (LLMs), reaso…
Echoing: Identity Failures when LLM Agents Talk to Each Other
Sarath Shekkizhar, Romain Cosentino, Adam Earle +1
As large language model (LLM) based agents interact autonomously with one another, a new class of failures emerges that cannot be predicted from single agent performance: behaviora…
Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
Bo Pang, Deqian Kong, Silvio Savarese +2
Reinforcement learning (RL) can elicit strong reasoning in large language models (LLMs), yet most open efforts focus on math and code. We propose Reasoning Curriculum, a simple two…
SCUBA: Salesforce Computer Use Benchmark
Yutong Dai, Krithika Ramakrishnan, Jing Gu +8
We introduce SCUBA, a benchmark designed to evaluate computer-use agents on customer relationship management (CRM) workflows within the Salesforce platform. SCUBA contains 300 task…