From the 1 of 13 linked papers with an AI index.
13 papers
J-CoT: Chain-of-Thought in J-Space
Junde Wu, Jiayuan Zhu, Fengling Liu +2
Chain-of-thought prompting improves language-model reasoning by carrying intermediate states across successive computation steps. However, relying on natural language as the only r…
Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming
Jiazhen Pan, Bailiang Jian, Paul Hager +19
The paper presents a dynamic red‑teaming framework (DAS) that continuously stress‑tests large language models on health tasks for robustness, privacy, bias, and hallucination, reve…
BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents
Jinge Wu, Hongjian Zhou, Mingde Zeng +8
Reproducing and comparing deep research agents today is hard: the same backbone evaluated on the same benchmark can report different accuracies across papers because the harness an…
From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding
Yuyuan Liu, Yiping Ji, Anjie Le +6
Finetuning Large Vision-Language Models with reinforcement learning has emerged as a promising approach to enhance their capability in object-level grounding. However, existing met…
MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows
Weixiang Shen, Chengzhi Shen, Yanzhu Hu +12
Medical imaging benchmarks often evaluate VLMs on pre-selected 2D images, slices, crops, or patches, making evaluation closer to visual recognition. Real clinical workflows impose…
Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded Reasoning
Jiayuan Zhu, Jiazhen Pan, Yuyuan Liu +2
The severe shortage of medical doctors limits access to timely and reliable healthcare, leaving millions underserved. Large language models (LLMs) offer a potential solution but st…