10 papers
Stress Testing Concept Erasure with Large Language Model Agents
Yuyang Xue, Feng Chen, Zhihua Liu +4
Concept erasure aims to remove semantic concepts from a trained generative model and is increasingly important for responsible AI deployment. However, verifying whether a model has…
CheXGenBench: A Unified Benchmark For Fidelity, Privacy and Utility of Synthetic Chest Radiographs
Raman Dutt, Pedro Sanchez, Yongchen Yao +3
Structured benchmarks have advanced text-conditional image generation for real-world imagery, however, no such benchmark exists for synthetic radiograph generation. Despite being a…
Why Do Vision Language Models Struggle To Recognize Human Emotions?
Madhav Agarwal, Sotirios A. Tsaftaris, Laura Sevilla-Lara +1
Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans. Vision-language models (VLMs) have made tremendous progress in the last…
A Causal Framework for Mitigating Data Shifts in Healthcare
Kurt Butler, Stephanie Riley, Damian Machlanski +13
Developing predictive models that perform reliably across diverse patient populations and heterogeneous environments is a core aim of medical research. However, generalization is o…
CSEval: A Framework for Evaluating Clinical Semantics in Text-to-Image Generation
Robert Cronshaw, Konstantinos Vilouras, Junyu Yan +4
Text-to-image generation has been increasingly applied in medical domains for various purposes such as data augmentation and education. Evaluating the quality and clinical reliabil…
Causal Ordering for Structure Learning from Time Series
Pedro P. Sanchez, Damian Machlanski, Steven McDonagh +1
Predicting causal structure from time series data is crucial for understanding complex phenomena in physiology, brain connectivity, climate dynamics, and socio-economic behaviour.…