From the 1 of 60 linked papers with an AI index.
2 citations · 5 across the 25 of their papers we have counts for
17 papers · 1 filter
Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking
Yepeng Huang, Jiawen Zhang, Michelle Dai +4
When a user question is underspecified, a capable model should recognize that its context is insufficient, identify the missing information, ask for it, and respond only once that…
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Xiaomin Li, Yuexing Hao, Jianheng Hou +90
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and inter…
An AI agent for treatment reasoning over a biomedical tool universe
Shanghua Gao, Ayush Noori, Richard Zhu +13
Treatment reasoning underpins every therapeutic decision, integrating disease context, comorbidities, medications, contraindications, and evolving biomedical knowledge to select an…
A global log for medical AI
Ayush Noori, Aaron E. Boussina, Hai Ho Bich +48
Modern computer systems rely on syslog, a universal protocol that records critical events across heterogeneous infrastructure. Medicine's rapidly growing AI stack has no equivalent…
Benchmarking AI Agents for Addressing Scientific Challenges Across Scales
Tianyu Liu, Allen Xin Wang, Antonia Panescu +30
AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in real research settings remain poorly understood. Existing benchma…
Evaluating Relational Reasoning in LLMs with REL
Lukas Fesser, Yasha Ektefaie, Ada Fang +2
Relational reasoning is the ability to infer relations that jointly bind multiple entities, attributes, or variables. This ability is central to scientific reasoning, but existing…