20 papers
VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus
Rohit Sinha, Kunal Tilaganji, Tanuja Ganu +3
Multimodal large language models often generate reasoning chains containing subtle errors that lead to incorrect answers. Current verification approaches have notable limitations.…
CPO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models
Swapnanil Mukherjee, Agyeya Negi, Tanuja Ganu +1
Current Multimodal Large Language Models (MLLMs) can process diverse sensory inputs, yet their reasoning remains heavily biased toward a dominant modality, resulting in brittle cro…
CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
Mercy Prasanna Ranjit, Anirban Porya, Sathvik Joel +8
A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive th…
Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models
Raj Jaiswal, Dhruv Jain, Rishabh Dhawan +4
Physics reasoning fails structurally in small language models: an error at any step propagates forward, corrupting every inference that follows. Limited domain knowledge, hallucina…
Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse
Raj Jaiswal, Anany Singh Divy, Savar Bhasin +3
Code language models are now trusted collaborators in production workflows for debugging, refactoring, and iterative repair, and every benchmark that evaluates them assumes the ins…
Cascaded Sparse Autoencoders Learn Multi-Level Visual Concepts in Multimodal LLMs
Yusong Zhao, Hengyi Wang, Tanuja Ganu +2
Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language tasks, yet their internal visual representations remain difficult to interpret. Spa…