2 papers
cs.AI2026
LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap
Yining Hua, Cyrus Ayubcha, Hongbin Na +4
Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimiz…
cs.AI2026
Design and Report Benchmarks for Knowledge Work
Yining Hua, Hongbin Na, Cyrus Ayubcha +1
The development of LLM agents has led to a growing body of work on knowledge-work AI, including coding, research, and healthcare. However, current knowledge-work evaluation and ben…