activity
20242026
most citedRethinking Retrieval-Augmented Generation for Medicine: A Large-Scale, Systematic Expert Evaluation and Practical Insights

1 citations · 1 across the 6 of their papers we have counts for

collaborators

11 papers

cs.CL2026

AgentGUI: An Interface for Observing and Steering Long-Running AI Agents

Xuan Zhao, Jiwoong Sohn, Qinyue Zheng +1

AI agents are increasingly adept at tackling complex, long-running tasks. With the rapid surge of autonomous capabilities, human oversight is systematically lagging behind due to l…

cs.AI2026

RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography

Mélanie Roschewitz, Kenneth Styppa, Yitian Tao +10

Vision-language models (VLM) have markedly advanced AI-driven interpretation and reporting of complex medical imaging, such as computed tomography (CT). Yet, existing methods large…

cs.AI2026

Process Reward Agents for Steering Knowledge-Intensive Reasoning

Jiwoong Sohn, Tomasz Sternal, Kenneth Styppa +2

Reasoning in knowledge-intensive domains remains challenging as intermediate steps are often not locally verifiable: unlike math or code, evaluating step correctness may require sy…

cs.CL2026

Med-CoReasoner: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning

Fan Gao, Sherry T. Tong, Jiwoong Sohn +11

While reasoning-enhanced large language models perform strongly on English medical tasks, a persistent multilingual gap remains, with substantially weaker reasoning in local langua…

cs.CL20251 cited

Rethinking Retrieval-Augmented Generation for Medicine: A Large-Scale, Systematic Expert Evaluation and Practical Insights

Hyunjae Kim, Jiwoong Sohn, Aidan Gilson +24

Large language models (LLMs) are transforming the landscape of medicine, yet two fundamental challenges persist: keeping up with rapidly evolving medical knowledge and providing ve…

cs.CL2025

Med-PRM: Medical Reasoning Models with Stepwise, Guideline-verified Process Rewards

Jaehoon Yun, Jiwoong Sohn, Jungwoo Park +9

Large language models have shown promise in clinical decision making, but current approaches struggle to localize and correct errors at specific steps of the reasoning process. Thi…