15 papers
Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking
Yepeng Huang, Jiawen Zhang, Michelle Dai +4
When a user question is underspecified, a capable model should recognize that its context is insufficient, identify the missing information, ask for it, and respond only once that…
An AI agent for treatment reasoning over a biomedical tool universe
Shanghua Gao, Ayush Noori, Richard Zhu +13
Treatment reasoning underpins every therapeutic decision, integrating disease context, comorbidities, medications, contraindications, and evolving biomedical knowledge to select an…
AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation
Shanghua Gao, Ada Fang, Marinka Zitnik
Scientific research proceeds through iterative cycles of hypothesis generation, experiment design, execution, and revision. AI agents can automate parts of this process, but existi…
When Sensors Fail: Temporal Sequence Models for Robust PPO under Sensor Drift
Kevin Vogt-Lowell, Theodoros Tsiligkaridis, Rodney Lafuente-Mercado +4
Real-world reinforcement learning systems must operate under distributional drift in their observation streams, yet most policy architectures implicitly assume fully observed and n…
Qworld: Question-Specific Evaluation Criteria for LLMs
Shanghua Gao, Yuchang Su, Pengwei Sui +2
Evaluating large language models (LLMs) on open-ended questions is difficult because response quality depends on the question's context. Binary scores and static rubrics fail to ca…
STRAND: Sequence-Conditioned Transport for Single-Cell Perturbations
Boyang Fu, George Dasoulas, Sameer Gabbita +5
Predicting how genetic perturbations change cellular state is a core problem for building controllable models of gene regulation. Perturbations targeting the same gene can produce…