3 papers
cs.CV2026
GroundBench: A Factorized, Counterfactual Benchmark for Locating VLM Affordance Failures
Sarthak Sattigeri
A companion evaluation found that naming the target part in a manipulation prompt increased action accuracy by 0.32-0.63 across eight vision-language models, with no model outperfo…
cs.CV2026
Part Grounding, Not Action Knowledge: Locating the Bottleneck in VLM Affordance Prediction
Sarthak Sattigeri
Benchmarks agree that vision-language models reason poorly about low-level manipulation, but an aggregate accuracy score does not say which step fails. We separate two steps that a…
cs.LG2026
Extending Beacon to Hindi: Cultural Adaptation Drives Cross-Lingual Sycophancy
Sarthak Sattigeri
Sycophancy, the tendency of language models to prioritize agreement with user preferences over principled reasoning, has been identified as a persistent alignment failure in Englis…