2 papers
cs.LG2025
Do Large Language Model Benchmarks Test Reliability?
Joshua Vendrow, Edward Vendrow, Sara Beery +1
When deploying large language models (LLMs), it is important to ensure that these models are not only capable, but also reliable. Many benchmarks have been created to track LLMs' g…
eess.IV2022
Understanding Transfer Learning for Chest Radiograph Clinical Report Generation with Modified Transformer Architectures
Edward Vendrow, Ethan Schonfeld
The image captioning task is increasingly prevalent in artificial intelligence applications for medicine. One important application is clinical report generation from chest radiogr…