most citedVerify when Uncertain: Beyond Self-Consistency in Black Box Hallucination Detection

1 citations · 1 across the 2 of their papers we have counts for

collaborators

10 papers

cs.CL20261 cited

Verify when Uncertain: Beyond Self-Consistency in Black Box Hallucination Detection

Yihao Xue, Kristjan Greenewald, Youssef Mroueh +1

Large Language Models (LLMs) often hallucinate, limiting their reliability in sensitive applications. In black-box settings, several self-consistency-based techniques have been pro…

cs.LG2026

Difference of Convex Programming in the Wasserstein Space with Applications to MMD Optimization

Clément Bonet, Pierre-Cyril Aubin-Frankowski, Youssef Mroueh

Optimizing functionals over the space of probability measures is now ubiquitous in machine learning. A widely used approach is to perform the optimization directly over the Wassers…

cs.LG2026

Guided Speculative Inference for Efficient Test-Time Alignment of LLMs

Jonathan Geuter, Youssef Mroueh, David Alvarez-Melis

We propose Guided Speculative Inference (GSI), a novel algorithm for efficient reward-guided decoding in large language models. GSI combines soft best-of- test-time scaling with…

cs.LG2025

GP-MoLFormer-Sim: Test Time Molecular Optimization through Contextual Similarity Guidance

Jiri Navratil, Jarret Ross, Payel Das +4

The ability to design molecules while preserving similarity to a target molecule and/or property is crucial for various applications in drug discovery, chemical design, and biology…

cs.LG2025

Reinforcement Learning with Verifiable Rewards: GRPO's Effective Loss, Dynamics, and Success Amplification

Youssef Mroueh

Group Relative Policy Optimization (GRPO) was introduced and used recently for promoting reasoning in LLMs under verifiable (binary) rewards. We show that the mean + variance calib…

cs.LG2025

KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity

Gholamali Aminian, Amir R. Asadi, Idan Shenfeld +1

Recent methods for aligning large language models (LLMs) with human feedback predominantly rely on a single reference model, which limits diversity, model overfitting, and underuti…