3 papers
cs.CR2026
CIPHER: Benchmarking Cross-record Inference over Privacy-Hardened Evidence Records
Suparno Roy Chowdhury, Manan Roy Choudhury, Dhruv Madhwal +1
Reasoning over privacy-constrained records requires combining structured attributes with evidence from free-text narratives. We introduce CIPHER (Cross-record Inference over Privac…
cs.LG2026
Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding
Eun Woo Im, Dhruv Madhwal, Vivek Gupta
Vision-Language Models demonstrate remarkable capabilities but often struggle with compositional reasoning, exhibiting vulnerabilities regarding word order and attribute binding. T…
cs.CL2026
Decomposed Prompting Does Not Fix Knowledge Gaps, But Helps Models Say "I Don't Know"
Dhruv Madhwal, Lyuxin David Zhang, Dan Roth +2
Large language models often struggle to recognize their knowledge limits in closed-book question answering, leading to confident hallucinations. While decomposed prompting is typic…