2 papers
cs.CL2026
Shattering the Shortcut: A Topology-Regularized Benchmark for Multi-hop Medical Reasoning in LLMs
Xing Zi, Xinying Zhou, Jinghao Xiao +2
While Large Language Models (LLMs) achieve expert-level performance on standard medical benchmarks through single-hop factual recall, they severely struggle with the complex, multi…
cs.CV2025
RSVLM-QA: A Benchmark Dataset for Remote Sensing Vision Language Model-based Question Answering
Xing Zi, Jinghao Xiao, Yunxiao Shi +4
Visual Question Answering (VQA) in remote sensing (RS) is pivotal for interpreting Earth observation data. However, existing RS VQA datasets are constrained by limitations in annot…