12 papers
One Model, Two Physical Stories: Auditing Misalignment in Multi-Modal World Modeling
Geigh Zollicoffer, Minh Vu, Rajiv Ranasinghe +1
World models, systems that generate what happens next given current environmental conditions, are increasingly being implemented with multi-modal generation in mind. However, gener…
Sanity Checks for Long-Form Hallucination Detection
Geigh Zollicoffer, Minh Vu, Hongli Zhan +2
Hallucination detection methods for large language models increasingly operate on chain-of-thought reasoning traces, yet it remains unclear whether they evaluate the reasoning itse…
World Model Robustness via Surprise Recognition
Geigh Zollicoffer, Tanush Chopra, Mingkuan Yan +3
AI systems deployed in the real world must contend with distractions and out-of-distribution (OOD) noise that can destabilize their policies and lead to unsafe behavior. While robu…
HalluField: Detecting LLM Hallucinations via Field-Theoretic Modeling
Minh Vu, Brian K. Tran, Syed A. Shah +3
Large Language Models (LLMs) exhibit impressive reasoning and question-answering capabilities. However, they often produce inaccurate or unreliable content known as hallucinations.…
MTRE: Multi-Token Reliability Estimation for Hallucination Detection in VLMs
Geigh Zollicoffer, Minh Vu, Manish Bhattarai
Vision-language models (VLMs) now rival human performance on many multimodal tasks, yet they still hallucinate objects or generate unsafe text. Current hallucination detectors, e.g…
Topological Signatures of Adversaries in Multimodal Alignments
Minh Vu, Geigh Zollicoffer, Huy Mai +3
Multimodal Machine Learning systems, particularly those aligning text and image data like CLIP/BLIP models, have become increasingly prevalent, yet remain susceptible to adversaria…