2 citations · 9 across the 34 of their papers we have counts for
35 papers
Knowing When Not to Answer: Abstention and Refusal Reasoning in Vision--Language Models
Karan Dua, Amit Agarwal, Hitesh Laxmichand Patel +8
Many medical conditions require diagnosis through detailed, multi-context clinical assessment rather than from visual appearance alone. Despite this, vision-language models (VLMs)…
No One Model Catches Every Harm: Benchmarking Content Moderation Across Safety Scenarios
Afshin Orojlooyjadid, Hitesh Patel
Large Language Models (LLMs) are increasingly deployed in real-world applications, yet they remain vulnerable to generating harmful content. From adversarial jailbreaks that bypass…
GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations
Jyotika Singh, Fang Tu, Aziza Mirsaidova +11
Benchmarks like GSM8K are popular measures of mathematical reasoning, but leaderboard gains can overstate true capability due to memorization of fixed test sets. Most robustness va…
Do Image-Text Metrics Respect Semantic Invariances?
Amit Agarwal, Hitesh Laxmichand Patel, Meizhu Liu +9
Reference-free image-to-text evaluators are now standard for scoring image-caption alignment, yet it is unclear whether they respect semantic invariances. We present an invariance…
Robust Audio-Text Retrieval via Cross-Modal Attention and Hybrid Loss
Meizhu Liu, Matthew Rowe, Amit Agarwal +8
Audio-text retrieval enables semantic alignment between audio content and natural language queries, supporting applications in multimedia search, accessibility, and surveillance. H…
SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks
Mohammadtaher Safarzadeh, Hitesh Laxmichand Patel, Afshin Orojlooyjadid +2
Large language models (LLMs) have achieved strong performance on natural language to SQL (NL2SQL) benchmarks, yet their reported accuracy may be inflated by contamination from benc…