activity
20242026
most citedMetaSynth: Meta-Prompting-Driven Agentic Scaffolds for Diverse Synthetic Data Generation

2 citations · 3 across the 9 of their papers we have counts for

collaborators

9 papers

cs.CV2026

Knowing When Not to Answer: Abstention and Refusal Reasoning in Vision--Language Models

Karan Dua, Amit Agarwal, Hitesh Laxmichand Patel +8

Many medical conditions require diagnosis through detailed, multi-context clinical assessment rather than from visual appearance alone. Despite this, vision-language models (VLMs)…

cs.CL2026

PLSQLBench: Benchmarking LLM Systems for Executable Procedural Database Programming

Marianne Menglin Liu, Leonid Boytsov, Daniel W. Peterson +13

We present PLSQLBench, to our knowledge the first benchmark for evaluating whether LLMs can write executable PL/SQL programs, with correctness measured through execution-based test…

cs.CL2026

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations

Jyotika Singh, Fang Tu, Aziza Mirsaidova +11

Benchmarks like GSM8K are popular measures of mathematical reasoning, but leaderboard gains can overstate true capability due to memorization of fixed test sets. Most robustness va…

cs.CL2026

SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks

Mohammadtaher Safarzadeh, Hitesh Laxmichand Patel, Afshin Orojlooyjadid +2

Large language models (LLMs) have achieved strong performance on natural language to SQL (NL2SQL) benchmarks, yet their reported accuracy may be inflated by contamination from benc…

cs.AI2025

FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models

Karan Dua, Hitesh Laxmichand Patel, Puneet Mittal +7

Developing document understanding models at enterprise scale requires large, diverse, and well-annotated datasets spanning a wide range of document types. However, collecting such…

cs.CL2025

Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge

Evangelia Spiliopoulou, Riccardo Fogliato, Hanna Burnsky +4

Large language models (LLMs) can serve as judges that offer rapid and reliable assessments of other LLM outputs. However, models may systematically assign overly favorable ratings…