7 papers
Aligning Language Model Benchmarks with Pairwise Preferences
Marco Gutierrez, Xinyi Leng, Hannah Cyberey +3
Language model benchmarks are pervasive and computationally-efficient proxies for real-world performance. However, many recent works find that benchmarks often fail to predict real…
Test-Time Hinting for Black-Box Vision-Language Models
Kaihua Hou, Abhijith Varma Mudunuri, Jiaxing Qiu +3
Test-time scaling (TTS) methods have proven highly effective for LLMs, yet their application to vision-language models (VLMs) remains relatively underexplored. Existing VLM TTS met…
ReasonEdit: Editing Vision-Language Models using Human Reasoning
Jiaxing Qiu, Kaihua Hou, Roxana Daneshjou +2
Model editing aims to correct errors in large, pretrained models without altering unrelated behaviors. While some recent works have edited vision-language models (VLMs), no existin…
Model Editing with Graph-Based External Memory
Yash Kumar Atri, Ahmed Alaa, Thomas Hartvigsen
Large language models (LLMs) have revolutionized natural language processing, yet their practical utility is often limited by persistent issues of hallucinations and outdated param…
A Systematic Analysis of Declining Medical Safety Messaging in Generative AI Models
Sonali Sharma, Ahmed M. Alaa, Roxana Daneshjou
Generative AI models, including large language models (LLMs) and vision-language models (VLMs), are increasingly used to interpret medical images and answer clinical questions. The…
Lifelong Knowledge Editing requires Better Regularization
Akshat Gupta, Phudish Prateepamornkul, Maochuan Lu +3
Knowledge editing is a promising way to improve factuality in large language models, but recent studies have shown significant model degradation during sequential editing. In this…