3 papers
cs.AI2026
Iterative Multimodal Retrieval-Augmented Generation for Medical Question Answering
Xupeng Chen, Binbin Shi, Chenqian Le +5
Medical retrieval-augmented generation (RAG) systems typically operate on text chunks extracted from biomedical literature, discarding the rich visual content (tables, figures, str…
cs.AI2026
Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models
Xupeng Chen, Binbin Shi, Chenqian Le +5
Vision-language models (VLMs) are increasingly applied to medical visual question answering (Med-VQA), yet whether they can \emph{localize} the evidence behind their answers---a pr…
cs.RO2025
RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning
Haoran Geng, Feishi Wang, Songlin Wei +34
Data scaling and standardized evaluation benchmarks have driven significant advances in natural language processing and computer vision. However, robotics faces unique challenges i…