Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap
Yining Hua, Cyrus Ayubcha, Hongbin Na +4
Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimiz…
cs.AI2026
Designing Benchmarks for Knowledge Work
Yining Hua, Hongbin Na, Cyrus Ayubcha +1
AI agents are moving quickly from answering isolated questions toward completing work through tools, software environments, and multi-step workflows. Much of what these systems are…