101 citations · 147 across the 36 of their papers we have counts for
12 papers · 1 filter
R2V Agent: Teaching SLMs When to Ask for Help
Raghu Vamshi Hemadri, Humaira Firdowse Mohammed, Rishabh Maheshwary +5
Efficient agentic systems should incur expensive frontier-model costs only on decisions where a cheaper local model is likely to fail. Existing LLM cascades usually route whole que…
Multi-scale Predictive Representations for Goal-conditioned Reinforcement Learning
Valliappan Chidambaram Adaikkappan, David Meger, Sai Rajeswar +1
This paper investigates robust representation learning in offline goal-conditioned reinforcement learning (GCRL). Particularly in sparse reward scenarios, learning representations…
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents
Xiangru Jian, Shravan Nayak, Kevin Qinghong Lin +5
Computer-use agents (CUAs) hold great promise for automating complex desktop workflows, yet progress toward general-purpose agents is bottlenecked by the scarcity of continuous, hi…
Grounding Computer Use Agents on Human Demonstrations
Aarash Feizi, Shravan Nayak, Xiangru Jian +14
Building reliable computer-use agents requires grounding: accurately connecting natural language instructions to the correct on-screen elements. While large datasets exist for web…
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
Aarash Feizi, Sai Rajeswar, Adriana Romero-Soriano +4
Understanding how effectively large vision language models (VLMs) compare visual inputs is crucial across numerous applications, yet this fundamental capability remains insufficien…
BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks
Juan Rodriguez, Xiangru Jian, Siba Smarak Panigrahi +40
Multimodal AI has the potential to significantly enhance document-understanding tasks, such as processing receipts, understanding workflows, extracting data from documents, and sum…