5 papers
StabilityBench: Benchmarking Instability in LLMs
Emma Kondrup, Zachary Yang, Anne Imouza +1
AI Assistants are increasingly deployed in high-stakes settings, such as healthcare or government services. Yet their real-world behavior remains poorly understood due to strong co…
CrediBench: Building Web-Scale Network Datasets for Information Integrity
Emma Kondrup, Sebastian Sabry, Hussein Abdallah +9
Automatically assessing the credibility of online sources presents an invaluable tool for navigating today's information ecosystem. However, existing approaches either depend on sc…
Dr. Bias: Social Disparities in AI-Powered Medical Guidance
Emma Kondrup, Anne Imouza
With the rapid progress of Large Language Models (LLMs), the general public now has easy and affordable access to applications capable of answering most health-related questions in…
Self-Exploring Language Models for Explainable Link Forecasting on Temporal Graphs via Reinforcement Learning
Zifeng Ding, Shenyang Huang, Zeyu Cao +11
Forecasting future links is a central task in temporal graph (TG) reasoning, requiring models to leverage historical interactions to predict upcoming ones. Traditional neural appro…
Are Large Language Models Good Temporal Graph Learners?
Shenyang Huang, Ali Parviz, Emma Kondrup +5
Large Language Models (LLMs) have recently driven significant advancements in Natural Language Processing and various other applications. While a broad range of literature has expl…