activity
20202026
most citedMeDAL: Medical Abbreviation Disambiguation Dataset for Natural Language Understanding Pretraining

25 citations · 53 across the 17 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG2026

Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection

Fatemeh Pesaran Zadeh, Seyeon Choi, Xing Han Lù +2

Large language models (LLMs) have enabled web agents that follow natural language goals through multi-step browser interactions. However, agents fine-tuned on specific trajectories…

cs.LG2026

Structured Distillation of Web Agent Capabilities Enables Generalization

Xing Han Lù, Siva Reddy

Frontier LLMs can navigate complex websites, but their cost and reliance on third-party APIs make local deployment impractical. We introduce Agent-as-Annotators, a framework that s…

cs.LG20251 cited

Grounding Computer Use Agents on Human Demonstrations

Aarash Feizi, Shravan Nayak, Xiangru Jian +14

Building reliable computer-use agents requires grounding: accurately connecting natural language instructions to the correct on-screen elements. While large datasets exist for web…

cs.LG20251 cited

Build the web for agents, not agents for the web

Xing Han Lù, Gaurav Kamath, Marius Mosbach +1

Recent advancements in Large Language Models (LLMs) and multimodal counterparts have spurred significant interest in developing web agents -- AI systems capable of autonomously nav…

cs.LG2025

AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories

Xing Han Lù, Amirhossein Kazemnejad, Nicholas Meade +7

Web agents enable users to perform tasks on web browsers through natural language interaction. Evaluating web agents trajectories is an important problem, since it helps us determi…

cs.LG2025

SafeArena: Evaluating the Safety of Autonomous Web Agents

Ada Defne Tur, Nicholas Meade, Xing Han Lù +6

LLM-based agents are becoming increasingly proficient at solving web-based tasks. With this capability comes a greater risk of misuse for malicious purposes, such as posting misinf…