most citedMedOmni-45°: A Safety-Performance Benchmark for Reasoning-Oriented LLMs in Medicine

1 citations · 1 across the 5 of their papers we have counts for

collaborators

8 papers

cs.CL2026

EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory

Ye Shen, Dun Pei, Yiqiu Guo +6

Despite recent advances in understanding and leveraging long-range conversational memory, existing benchmarks still lack systematic evaluation of large language models(LLMs) across…

cs.CL2025

Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs

Junying Wang, Zicheng Zhang, Ye Shen +8

High-quality, multi-modal benchmarks are crucial for advancing scientific reasoning in large models yet their manual creation is costly and unscalable. To address this bottleneck,…

cs.CL2025

A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation

Ye Shen, Junying Wang, Farong Wen +4

The rapid progress of Multi-Modal Large Language Models (MLLMs) has spurred the creation of numerous benchmarks. However, conventional full-coverage Question-Answering evaluations…

cs.CV20251 cited

MedOmni-45°: A Safety-Performance Benchmark for Reasoning-Oriented LLMs in Medicine

Kaiyuan Ji, Yijin Guo, Zicheng Zhang +4

With the increasing use of large language models (LLMs) in medical decision-support, it is essential to evaluate not only their final answers but also the reliability of their reas…

cs.CL2025

User-centric Subjective Leaderboard by Customizable Reward Modeling

Qi Jia, Xiujie Song, Zicheng Zhang +4

Existing benchmarks for large language models (LLMs) predominantely focus on assessing their capabilities through verifiable tasks. Such objective and static benchmarks offer limit…

cs.CL2025

Affordance Benchmark for MLLMs

Junying Wang, Wenzhe Li, Yalun Wu +6

Affordance theory suggests that environments inherently provide action possibilities shaping perception and behavior. While Multimodal Large Language Models (MLLMs) achieve strong…