activity
20242026
collaborators

11 papers

cs.LG2026

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives

Matthieu Bou, Nyal Patel, Arjun Jagota +2

The objectives that Large Language Models (LLMs) implicitly optimize remain dangerously opaque, making trustworthy alignment and auditing a grand challenge. While Inverse Reinforce…

cs.AI2026

Why LLMs Fail at Causal Discovery and How Interventional Agents Escape

Amartya Roy, Sonali Parbhoo

Causal discovery is a cornerstone of scientific reasoning, yet whether large language models can perform it reliably remains an open question. Recent benchmarks show that even fine…

cs.LG2026

Causal methods for LLM development and evaluation

Dennis Frauen, Marie Brockschmidt, Konstantin Hess +10

Large language model (LLM) development is currently driven by large-scale empirical iteration over data mixtures, reward models, routing strategies, and evaluation pipelines. Here,…

cs.LG2026

Causal Machine Learning Is Not a Panacea: A Roadmap for Observational Causal Inference in Health

Donna Tjandra, Trenton Chang, Sonali Parbhoo +8

Objective: The growing availability of large-scale observational clinical datasets and challenges in conducting randomized controlled trials have spurred enthusiasm in using causal…

cs.LG2026

Learning from Failures: Understanding LLM Alignment through Failure-Aware Inverse RL

Nyal Patel, Matthieu Bou, Arjun Jagota +2

Reinforcement Learning from Human Feedback (RLHF) aligns Large Language Models (LLMs) with human preferences, yet the underlying reward signals they internalize remain hidden, posi…

cs.CL2025

Insights from the Inverse: Reconstructing LLM Training Goals Through Inverse Reinforcement Learning

Jared Joselowitz, Ritam Majumdar, Arjun Jagota +4

Large language models (LLMs) trained with Reinforcement Learning from Human Feedback (RLHF) have demonstrated remarkable capabilities, but their underlying reward functions and dec…