10 papers
Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur?
Apratim Bhattacharyya, Shweta Mahajan, Sanjay Haresh +5
Learning everyday skills, like cooking a dish, relies increasingly on instructional media such as online videos. This opens the door to the use of video (and multimodal) large lang…
From Out-of-Distribution Detection to Hallucination Detection: A Geometric View
Litian Liu, Reza Pourreza, Yubing Jian +2
Detecting hallucinations in large language models is a critical open problem with significant implications for safety and reliability. While existing hallucination detection method…
RoCA: Robust Cross-Domain End-to-End Autonomous Driving
Rajeev Yasarla, Shizhong Han, Hsin-Pai Cheng +7
End-to-end (E2E) autonomous driving has recently emerged as a new paradigm, offering significant potential. However, few studies have looked into the practical challenge of deploym…
Enhancing Hallucination Detection through Noise Injection
Litian Liu, Reza Pourreza, Sunny Panchal +4
Large Language Models (LLMs) are prone to generating plausible yet incorrect responses, known as hallucinations. Effectively detecting hallucinations is therefore crucial for the s…
MAPLE: Latent Multi-Agent Play for End-to-End Autonomous Driving
Rajeev Yasarla, Deepti Hegde, Hsin-Pai Cheng +9
Vision-language-action (VLA) models are effective as end-to-end motion planners, but can be brittle when evaluated in closed-loop settings due to being trained under traditional im…
Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
Apratim Bhattacharyya, Bicheng Xu, Sanjay Haresh +6
Multi-modal Large Language Models (LLM) have advanced conversational abilities but struggle with providing live, interactive step-by-step guidance, a key capability for future AI a…