activity
20242026
collaborators

6 papers

cs.AI2026

SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy

Peiyao Xiao, Xiaogang Li, Xinyi Gao +7

As LLMs achieved breakthroughs in general reasoning, their proficiency in specialized scientific domains reveals pronounced gaps in existing benchmarks due to data contamination, i…

cs.CV2026

Medical thinking with multiple images

Zonghai Yao, Benlu Wang, Yifan Zhang +8

Large language models perform well on many medical QA benchmarks, but real clinical reasoning often requires integrating evidence across multiple images rather than interpreting a…

cs.AI2026

Rethinking Patient Education as Multi-turn Multi-modal Interaction

Zonghai Yao, Zhipeng Tang, Chengtao Lin +5

Most medical multimodal benchmarks focus on static tasks such as image question answering, report generation, and plain-language rewriting. Patient education is more demanding: sys…

cs.CL2026

From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations

Benlu Wang, Iris Xia, Yifan Zhang +6

Large language models (LLMs) have demonstrated promising performance on medical benchmarks; however, their ability to perform medical calculations, a crucial aspect of clinical dec…

cs.CL2025

Balancing Knowledge Delivery and Emotional Comfort in Healthcare Conversational Systems

Shang-Chi Tsai, Yun-Nung Chen

With the advancement of large language models, many dialogue systems are now capable of providing reasonable and informative responses to patients' medical conditions. However, whe…

cs.CL2024

GPT-4o System Card

OpenAI, :, Aaron Hurst +416

GPT-4o is an autoregressive omni model that accepts as input any combination of text, audio, image, and video, and generates any combination of text, audio, and image outputs. It's…