Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Xiaomin Li, Yuexing Hao, Jianheng Hou +90
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and inter…
cs.AI2026
Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents
Muyu He, Anand Kumar, Tsach Mackey +3
Despite rapid progress in building conversational AI agents, robustness is still largely untested. Small shifts in user behavior, such as being more impatient, incoherent, or skept…
cs.AI2026
ReasonOps: Operator Segmentation for LLM Reasoning Traces
Daniel Lee, Owen Queen, James Zou
Chain-of-thought traces from large reasoning models can span tens of thousands of tokens, yet we lack a vocabulary for describing their internal structure. Previous methods develop…