19 papers
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
Rui Yang, Weihao Xuan, Yi Lin +23
Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive discl…
Robust Self-Supervised Cross-Modal Super-Resolution against Real-World Misaligned Observations
Xiaoyu Dong, Jiahuan Li, Ziteng Cui +1
Cross-modal super-resolution (SR) on real-world misaligned data is challenging, as only unlabeled low-resolution (LR) source and high-resolution (HR) guide images with complex spat…
MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos
Leyuan Yu, Xiao Tang, Minghao Liu +6
Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input. Existing what-if tasks typically v…
Beyond Monolingual Deep Research: Evaluating Agents and Retrievers with Cross-Lingual BrowseComp-Plus
Yuheng Lu, Qingcheng Zeng, Heli Qi +6
Deep research agents are increasingly evaluated on their ability to search for evidence, reason over retrieved sources, and produce grounded answers. Existing browsing benchmarks,…
Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards
Fang Wu, Aaron Tu, Weihao Xuan +21
Reinforcement learning with verifiable rewards (RLVR) is a practical, scalable way to improve large language models on math, code, and other structured tasks. However, we argue tha…
Proteo-R1: Reasoning Foundation Models for De Novo Protein Design
Fang Wu, Weihao Xuan, Heli Qi +26
Deep learning in de novo protein design has achieved atomic-level fidelity. However, existing models remain largely non-deliberative: they directly synthesize molecular geometries…