8 papers
Learning complete and explainable visual representations from itemized text supervision
Yiwei Lyu, Chenhui Zhao, Soumyanil Banerjee +5
Training vision models with language supervision enables general and transferable representations. However, many visual domains, especially non-object-centric domains such as medic…
CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization
Xinhai Hou, Shaoyuan Xu, Manan Biyani +4
Agentic vision-language models are increasingly trained to "think with images" by calling image operations. However, we show that high final-answer accuracy often hides unfaithful…
Towards Scalable Language-Image Pre-training for 3D Medical Imaging
Chenhui Zhao, Yiwei Lyu, Asadur Chowdury +6
The scalability of current language-image pre-training for 3D medical imaging, such as CT and MRI, is constrained by the need for radiologists to manually curate raw clinical studi…
Learning neuroimaging models from health system-scale data
Yiwei Lyu, Samir Harake, Asadur Chowdury +18
Neuroimaging is a ubiquitous tool for evaluating patients with neurological diseases. The global demand for magnetic resonance imaging (MRI) studies has risen steadily, placing sig…
CNS-Obsidian: A Neurosurgical Vision-Language Model Built From Scientific Publications
Anton Alyakin, Jaden Stryker, Daniel Alexander Alber +29
General-purpose VLMs demonstrate impressive capabilities, but their opaque training on uncurated internet data poses critical limitations for high-stakes decision-making, such as i…
Health system learning achieves generalist neuroimaging models
Akhil Kondepudi, Akshay Rao, Chenhui Zhao +14
Frontier artificial intelligence (AI) models, such as OpenAI's GPT-5 and Meta's DINOv3, have advanced rapidly through training on internet-scale public data, yet such systems lack…