collaborators

5 papers

cs.LG2026

The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems

Richard Ren, Arunim Agarwal, Mantas Mazeika +13

As large language models (LLMs) become more capable and agentic, the requirement for trust in their outputs grows significantly, yet at the same time concerns have been mounting th…

cs.CL2025

Synthetic Dataset for Evaluating Complex Compositional Knowledge for Natural Language Inference

Sushma Anand Akoju, Robert Vacareanu, Haris Riaz +2

We introduce a synthetic dataset called Sentences Involving Complex Compositional Knowledge (SICCK) and a novel analysis that investigates the performance of Natural Language Infer…

cs.CL2025

Online Rubrics Elicitation from Pairwise Comparisons

MohammadHossein Rezaei, Robert Vacareanu, Zihao Wang +4

Rubrics provide a flexible way to train LLMs on open-ended long-form answers where verifiable rewards are not applicable and human preferences provide coarse signals. Prior work sh…

cs.CL2025

Jailbreaking to Jailbreak

Jeremy Kritz, Vaughn Robinson, Robert Vacareanu +7

Large Language Models (LLMs) can be used to red team other models (e.g. jailbreaking) to elicit harmful contents. While prior works commonly employ open-weight models or private un…

cs.CL2025

MorphNLI: A Stepwise Approach to Natural Language Inference Using Text Morphing

Vlad Andrei Negru, Robert Vacareanu, Camelia Lemnaru +2

We introduce MorphNLI, a modular step-by-step approach to natural language inference (NLI). When classifying the premise-hypothesis pairs into {entailment, contradiction, neutral},…