activity
20242026
collaborators

7 papers

cs.LG2026

Large Language Models Should Learn Personalized Rather Than Aggregated Human Preferences

Cristina Garbacea

Current approaches to aligning large language models (LLMs) aggregate diverse human preferences into a single reward signal, effectively optimizing for a hypothetical ``average use…

cs.AI2026

Personalized Benchmarking: Evaluating LLMs by Individual Preferences

Cristina Garbacea, Heran Wang, Chenhao Tan

With the rise in capabilities of large language models (LLMs) and their deployment in real-world tasks, evaluating LLM alignment with human preferences has become an important chal…

cs.CL2025

HyPerAlign: Interpretable Personalized LLM Alignment via Hypothesis Generation

Cristina Garbacea, Chenhao Tan

Alignment algorithms are widely used to align large language models (LLMs) to human users based on preference annotations. Typically these (often divergent) preferences are aggrega…

cs.CL2025

RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals

David Reber, Sean Richardson, Todd Nief +2

Reward models are widely used as proxies for human preferences when aligning or evaluating LLMs. However, reward models are black boxes, and it is often unclear what, exactly, they…

cs.AI2025

Evaluating the Goal-Directedness of Large Language Models

Tom Everitt, Cristina Garbacea, Alexis Bellot +4

To what extent do LLMs use their capabilities towards their given goal? We take this as a measure of their goal-directedness. We evaluate goal-directedness on tasks that require in…

cs.CL2025

Why is constrained neural language generation particularly challenging?

Cristina Garbacea, Qiaozhu Mei

Recent advances in deep neural language models combined with the capacity of large scale datasets have accelerated the development of natural language generation systems that produ…