collaborators

16 papers

cs.CL2026

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases

Rui Yang, Weihao Xuan, Yi Lin +23

Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive discl…

cs.CL2026

Beyond Monolingual Deep Research: Evaluating Agents and Retrievers with Cross-Lingual BrowseComp-Plus

Yuheng Lu, Qingcheng Zeng, Heli Qi +6

Deep research agents are increasingly evaluated on their ability to search for evidence, reason over retrieved sources, and produce grounded answers. Existing browsing benchmarks,…

cs.AI2026

Knowledge Index of Noah's Ark

Sheng Jin, Minghao Liu, Yunze Xiao +24

Knowledge benchmarks for LLMs face three issues: scaling-driven designs that do not operationalize disciplinary representativeness; flat-payment annotation that permits lazy consen…

cs.LG2026

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards

Fang Wu, Aaron Tu, Weihao Xuan +21

Reinforcement learning with verifiable rewards (RLVR) is a practical, scalable way to improve large language models on math, code, and other structured tasks. However, we argue tha…

cs.AI2026

Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations

Junjue Wang, Weihao Xuan, Heli Qi +7

Operational disaster response goes beyond damage assessment, requiring responders to integrate multi-sensor signals, reason over road networks, populations and key facilities, plan…

cs.LG2026

Proteo-R1: Reasoning Foundation Models for De Novo Protein Design

Fang Wu, Weihao Xuan, Heli Qi +26

Deep learning in de novo protein design has achieved atomic-level fidelity. However, existing models remain largely non-deliberative: they directly synthesize molecular geometries…