45 papers
Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking
Yepeng Huang, Jiawen Zhang, Michelle Dai +4
When a user question is underspecified, a capable model should recognize that its context is insufficient, identify the missing information, ask for it, and respond only once that…
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Xiaomin Li, Yuexing Hao, Jianheng Hou +90
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and inter…
Kidney function and kidney failure prediction in a large multiethnic population
James A. Diao, Morgan Sanchez, Adir Sommer +7
Background: Patients with chronic kidney disease (CKD) experience worsening kidney function and develop subsequent kidney failure at different rates. Accurate estimates of a patien…
Adaptive Time Series Reasoning via Segment Selection
Shvat Messica, Jiawen Zhang, Kevin Li +2
The paper presents ARTIST, a model that adaptively selects informative temporal segments of a time series to answer natural‑language questions, using a controller‑reasoner architec…
How Post-Training Shapes Biological Reasoning Models
Lukas Fesser, Hanlin Zhang, Michelle M. Li +5
Scientific reasoning models for biology combine language models with foundation models trained on multimodal biological data, including DNA, RNA, and proteins. These models are bui…
An AI agent for treatment reasoning over a biomedical tool universe
Shanghua Gao, Ayush Noori, Richard Zhu +13
Treatment reasoning underpins every therapeutic decision, integrating disease context, comorbidities, medications, contraindications, and evolving biomedical knowledge to select an…