most citedVerify when Uncertain: Beyond Self-Consistency in Black Box Hallucination Detection

1 citations · 1 across the 1 of their papers we have counts for

collaborators

5 papers

cs.CL20261 cited

Verify when Uncertain: Beyond Self-Consistency in Black Box Hallucination Detection

Yihao Xue, Kristjan Greenewald, Youssef Mroueh +1

Large Language Models (LLMs) often hallucinate, limiting their reliability in sensitive applications. In black-box settings, several self-consistency-based techniques have been pro…

cs.LG2026

Beyond What Seems Necessary: Hidden Gains from Scaling Training-Time Reasoning Length under Outcome Supervision

Yihao Xue, Allan Zhang, Jianhao Huang +2

Training LLMs to think and reason for longer has become a key ingredient in building state-of-the-art models that can solve complex problems previously out of reach. Recent efforts…

cs.AI2026

LoRA is All You Need for Safety Alignment of Reasoning LLMs

Yihao Xue, Baharan Mirzasoleiman

Reasoning-capable LLMs have achieved major breakthroughs in solving complex problems, but recent work shows that acquiring and deploying strong reasoning can introduce significant…

cs.LG2025

Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions

Yihao Xue, Jiping Li, Baharan Mirzasoleiman

Weak-to-Strong Generalization (W2SG), where a weak model supervises a stronger one, serves as an important analogy for understanding how humans might guide superhuman intelligence…

cs.LG2025

Challenges and Opportunities in Improving Worst-Group Generalization in Presence of Spurious Features

Siddharth Joshi, Yu Yang, Yihao Xue +2

Deep neural networks often exploit *spurious* features that are present in the majority of examples within a class during training. This leads to *poor worst-group test accuracy*,…