2 citations · 2 across the 4 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Strong Reasoning Isn't Enough: Evaluating Evidence Elicitation in Interactive Diagnosis
Zhuohan Long, Zhijie Bao, Zhongyu Wei
Interactive medical consultation requires an agent to proactively elicit missing clinical evidence under uncertainty. Yet existing evaluations largely remain static or outcome-cent…
cs.CL2024
From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking
Siyuan Wang, Zhuohan Long, Zhihao Fan +1
The rapid development of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) has exposed vulnerabilities to various adversarial attacks. This paper provides a…
cs.CL2024★ 2 cited
Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation
Siyuan Wang, Zhuohan Long, Zhihao Fan +2
This paper presents a benchmark self-evolving framework to dynamically evaluate rapidly advancing Large Language Models (LLMs), aiming for a more accurate assessment of their capab…