activity
20242026
most citedMedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models

1 citations · 1 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CV2026

ChartE: A Comprehensive Benchmark for End-to-End Chart Editing

Shuo Li, Jiajun Sun, Zhekai Wang +9

Charts are a fundamental visualization format for structured data analysis. Enabling end-to-end chart editing according to user intent is of great practical value, yet remains chal…

cs.AI2026

RouteMoA: Dynamic Routing without Pre-Inference Boosts Efficient Mixture-of-Agents

Jize Wang, Han Wu, Zhiyuan You +9

Mixture-of-Agents (MoA) improves LLM performance through layered collaboration, but its dense topology raises costs and latency. Existing methods employ LLM judges to filter respon…

cs.CL20251 cited

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models

Xiaomin Li, Mingye Gao, Yuexing Hao +4

Clinical guidelines, typically structured as decision trees, are central to evidence-based medical practice and critical for ensuring safe and accurate diagnostic decision-making.…

cs.CV2025

On the Perception Bottleneck of VLMs for Chart Understanding

Junteng Liu, Weihao Zeng, Xiwen Zhang +3

Chart understanding requires models to effectively analyze and reason about numerical data, textual elements, and complex visual components. Our observations reveal that the percep…

cs.AI2024

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners

Weihao Zeng, Yuzhen Huang, Lulu Zhao +3

In the absence of extensive human-annotated data for complex reasoning tasks, self-improvement -- where models are trained on their own outputs -- has emerged as a primary method f…