activity
20242026
collaborators

15 papers

cs.AI2026

Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking

Yepeng Huang, Jiawen Zhang, Michelle Dai +4

When a user question is underspecified, a capable model should recognize that its context is insufficient, identify the missing information, ask for it, and respond only once that…

cs.AI2026

An AI agent for treatment reasoning over a biomedical tool universe

Shanghua Gao, Ayush Noori, Richard Zhu +13

Treatment reasoning underpins every therapeutic decision, integrating disease context, comorbidities, medications, contraindications, and evolving biomedical knowledge to select an…

cs.AI2026

AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation

Shanghua Gao, Ada Fang, Marinka Zitnik

Scientific research proceeds through iterative cycles of hypothesis generation, experiment design, execution, and revision. AI agents can automate parts of this process, but existi…

cs.LG2026

When Sensors Fail: Temporal Sequence Models for Robust PPO under Sensor Drift

Kevin Vogt-Lowell, Theodoros Tsiligkaridis, Rodney Lafuente-Mercado +4

Real-world reinforcement learning systems must operate under distributional drift in their observation streams, yet most policy architectures implicitly assume fully observed and n…

cs.CL2026

Qworld: Question-Specific Evaluation Criteria for LLMs

Shanghua Gao, Yuchang Su, Pengwei Sui +2

Evaluating large language models (LLMs) on open-ended questions is difficult because response quality depends on the question's context. Binary scores and static rubrics fail to ca…

q-bio.GN2026

STRAND: Sequence-Conditioned Transport for Single-Cell Perturbations

Boyang Fu, George Dasoulas, Sameer Gabbita +5

Predicting how genetic perturbations change cellular state is a core problem for building controllable models of gene regulation. Perturbations targeting the same gene can produce…