From the 1 of 5 linked papers with an AI index.
5 papers
AgentGUI: An Interface for Observing and Steering Long-Running AI Agents
Xuan Zhao, Jiwoong Sohn, Qinyue Zheng +1
AgentGUI is a locally hosted graphical interface that lets users observe and steer AI agents during long-running tasks, providing trajectory visualizations, manual and automated st…
Process Reward Agents for Steering Knowledge-Intensive Reasoning
Jiwoong Sohn, Tomasz Sternal, Kenneth Styppa +2
Reasoning in knowledge-intensive domains remains challenging as intermediate steps are often not locally verifiable: unlike math or code, evaluating step correctness may require sy…
MTBBench: A Multimodal Sequential Clinical Decision-Making Benchmark in Oncology
Kiril Vasilev, Alexandre Misrahi, Eeshaan Jain +5
Multimodal Large Language Models (LLMs) hold promise for biomedical reasoning, but current benchmarks fail to capture the complexity of real-world clinical workflows. Existing eval…
AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
Samuel Schmidgall, Rojin Ziaei, Carl Harris +3
Evaluating large language models (LLM) in clinical scenarios is crucial to assessing their potential clinical utility. Existing benchmarks rely heavily on static question-answering…
AgentRxiv: Towards Collaborative Autonomous Research
Samuel Schmidgall, Michael Moor
Progress in scientific discovery is rarely the result of a single "Eureka" moment, but is rather the product of hundreds of scientists incrementally working together toward a commo…