4 papers
Does Playing it Safe Count as Faithfulness? Reassessing LVLM Hallucination Mitigation Methods
Mehrdad Fazli, Sina Mansouri, Mohit Marvania +1
Recent inference-time hallucination mitigation methods for large vision-language models (LVLMs) report strong gains on hallucination benchmarks. However, it remains unclear whether…
Stability-Aware Feature Design for Robust Watermark Detection in Machine-Generated Text
Sina Mansouri, Mohit Marvania, Abolfazl Safikhani
The widespread adoption of large language models (LLMs) has intensified the demand for principled methods to distinguish human from machine-generated text. Watermarking provides a…
VeriSim: A Configurable Framework for Stress-Testing Medical AI Under Patient Communication Noise
Sina Mansouri, Mohit Marvania, Vibhavari Ashok Shihorkar +5
Medical large language models are typically evaluated on idealized patient cases that do not reflect how real patients communicate. We introduce VeriSim, a patient simulation frame…
Batch Prompting Suppresses Overthinking Reasoning Under Constraint: How Batch Prompting Suppresses Overthinking in Reasoning Models
Saurabh Srivastava, Janit Bidhan, Hao Yan +7
Large Reasoning Models (LRMs) achieve strong performance through explicit chain-of-thought reasoning but suffer from \textit{overthinking}: generating excessive reasoning tokens ev…