3 papers
cs.LG2026
Grounding Computer Use Agents on Human Demonstrations
Aarash Feizi, Shravan Nayak, Xiangru Jian +14
Building reliable computer-use agents requires grounding: accurately connecting natural language instructions to the correct on-screen elements. While large datasets exist for web…
cs.LG2025
Weak Supervision for Real World Graphs
Pratheeksha Nair, Reihaneh Rabbany
Node classification in real world graphs often suffers from label scarcity and noise, especially in high stakes domains like human trafficking detection and misinformation monitori…
cs.LG2025
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
Aarash Feizi, Sai Rajeswar, Adriana Romero-Soriano +4
Understanding how effectively large vision language models (VLMs) compare visual inputs is crucial across numerous applications, yet this fundamental capability remains insufficien…