1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CL2024★ 1 cited
Revisiting the Superficial Alignment Hypothesis
Mohit Raghavendra, Vaskar Nath, Sean Hendryx
The Superficial Alignment Hypothesis posits that almost all of a language model's abilities and knowledge are learned during pre-training, while post-training is about giving a mod…
cs.CL2024
Pre-Training Multimodal Hallucination Detectors with Corrupted Grounding Data
Spencer Whitehead, Jacob Phillips, Sean Hendryx
Multimodal language models can exhibit hallucinations in their outputs, which limits their reliability. The ability to automatically detect these errors is important for mitigating…