works on

From the 1 of 5 linked papers with an AI index.

most citedAddressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

2 citations · 2 across the 1 of their papers we have counts for

collaborators

5 papers

cs.LG20262 cited

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

Jiazhen Pan, Bailiang Jian, Paul Hager +19

The paper presents a dynamic red‑teaming framework (DAS) that continuously stress‑tests large language models on health tasks for robustness, privacy, bias, and hallucination, reve…

cs.AI2026

RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation

Chengzhi Shen, Weixiang Shen, Tobias Susetzky +7

Intensive care units (ICU) generate long, dense and evolving streams of clinical information, where physicians must repeatedly reassess patient states under time pressure, undersco…

cs.SE2026

Git Context Controller: Manage the Context of LLM-based Agents like Git

Junde Wu, Minhao Hu, Jiayuan Zhu +4

Large language model (LLM) agents have demonstrated strong capabilities in long-horizon tasks by interleaving reasoning with tool use. However, as these agents scale to complex wor…

cs.LG2026

Evo: Autoregressive-Diffusion Large Language Models with Evolving Balance

Junde Wu, Minhao Hu, Jiayuan Zhu +7

We introduce \textbf{Evo}, a duality latent trajectory model that bridges autoregressive (AR) and diffusion-based language generation within a continuous evolutionary generative fr…

cs.AI2025

Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools

Junde Wu, Jiayuan Zhu, Yuyuan Liu +2

We introduce Agentic Reasoning, a framework that enhances large language model (LLM) reasoning by integrating external tool-using agents. Agentic Reasoning dynamically leverages we…