activity
20212026
most citedMind2Web: Towards a Generalist Agent for the Web

30 citations · 78 across the 17 of their papers we have counts for

collaborators
Showing 2024Show all

6 papers · 1 filter

cs.AI2024

Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents

Yu Gu, Kai Zhang, Yuting Ning +9

Language agents based on large language models (LLMs) have demonstrated great promise in automating web-based tasks. Recent work has shown that incorporating advanced planning algo…

cs.AI2024★ 3 cited

Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Boyu Gou, Ruohan Wang, Boyuan Zheng +5

Multimodal large language models (MLLMs) are transforming the capabilities of graphical user interface (GUI) agents, facilitating their transition from controlled simulations to co…

cs.CV2024★ 1 cited

Dual-View Visual Contextualization for Web Navigation

Jihyung Kil, Chan Hee Song, Boyuan Zheng +3

Automatic web navigation aims to build a web agent that can follow language instructions to execute complex and diverse tasks on real-world websites. Existing work primarily takes…

cs.CL2024

A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents

Lingbo Mo, Zeyi Liao, Boyuan Zheng +3

Language agents powered by large language models (LLMs) have seen exploding development. Their capability of using language as a vehicle for thought and communication lends an incr…

cs.CL2024★ 3 cited

The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts

Lingfeng Shen, Weiting Tan, Sihao Chen +6

As the influence of large language models (LLMs) spans across global communities, their safety challenges in multilingual settings become paramount for alignment research. This pap…

cs.IR2024★ 10 cited

GPT-4V(ision) is a Generalist Web Agent, if Grounded

Boyuan Zheng, Boyu Gou, Jihyung Kil +2

The recent development on large multimodal models (LMMs), especially GPT-4V(ision) and Gemini, has been quickly expanding the capability boundaries of multimodal models beyond trad…