5 papers
Justified or Just Convincing? Error Verifiability as a Dimension of LLM Quality
Xiaoyuan Zhu, Kimberly Le Truong, Riccardo Fogliato +8
As LLMs are deployed in high-stakes settings, users must judge the correctness of individual responses, often relying on model-generated justifications such as reasoning chains or…
M^4olGen: Multi-Agent, Multi-Stage Molecular Generation under Precise Multi-Property Constraints
Yizhan Li, Florence Cloutier, Sifan Wu +5
Generating molecules that satisfy precise numeric constraints over multiple physicochemical properties is critical and challenging. Although large language models (LLMs) are expres…
What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles
Mengtao Zhou, Sifan Wu, Huan Zhang +2
We investigate the capacity of Large Language Models (LLMs) for imaginative reasoning--the proactive construction, testing, and revision of hypotheses in information-sparse environ…
Improving Clinical Note Generation from Complex Doctor-Patient Conversation
Yizhan Li, Sifan Wu, Christopher Smith +2
Writing clinical notes and documenting medical exams is a critical task for healthcare professionals, serving as a vital component of patient care documentation. However, manually…
Seeing Beyond Words: MatVQA for Challenging Visual-Scientific Reasoning in Materials Science
Sifan Wu, Huan Zhang, Yizhan Li +3
The emergence of Multimodal Large Language Models (MLLMs) that integrate vision and language modalities has unlocked new potentials for scientific reasoning, outperforming prior be…