8 papers
ResidencyRL: Reinforcement Learning in Simulated Clinical Environments
Valentin Liévin, Samuel Schmidgall, Tim Strother +32
In medical education, physicians convert academic knowledge into clinical expertise through residency: years of training across thousands of encounters, with diverse sources of fee…
Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual Illusions
Xiaoxiao Sun, Mingyang Li, Kun Yuan +7
Large Vision-Language Models (VLMs) often answer classic visual illusions "correctly" on original images, yet persist with the same responses when illusion factors are inverted, ev…
From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature
Kun Yuan, Min Woo Sun, Zhen Chen +7
There is a growing interest in developing strong biomedical vision-language models. A popular approach to achieve robust representations is to use web-scale scientific data. Howeve…
Beyond Static Visual Tokens: Structured Sequential Visual Chain-of-Thought Reasoning
Guangfu Guo, Xiaoqian Lu, Yue Feng +1
Current multimodal LLMs encode images as static visual prefixes and rely on text-based reasoning, lacking goal-driven and adaptive visual access. Inspired by human visual perceptio…
PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
James Burgess, Jan N. Hansen, Duo Peng +5
Search agents are language models (LMs) that reason and search knowledge bases (or the web) to answer questions; recent methods supervise only the final answer accuracy using reinf…
No Tokens Wasted: Leveraging Long Context in Biomedical Vision-Language Models
Min Woo Sun, Alejandro Lozano, Javier Gamazo Tejero +8
Embedding vision-language models (VLMs) are typically pretrained with short text windows (<77 tokens), which forces the truncation of long-format captions. Yet, the distribution of…