101 citations · 161 across the 40 of their papers we have counts for
Showing 2025 · cs.LGShow all
2 papers · 2 filters
cs.LG2025★ 1 cited
Grounding Computer Use Agents on Human Demonstrations
Aarash Feizi, Shravan Nayak, Xiangru Jian +14
Building reliable computer-use agents requires grounding: accurately connecting natural language instructions to the correct on-screen elements. While large datasets exist for web…
cs.LG2025
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
Aarash Feizi, Sai Rajeswar, Adriana Romero-Soriano +4
Understanding how effectively large vision language models (VLMs) compare visual inputs is crucial across numerous applications, yet this fundamental capability remains insufficien…