works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.CV2026

MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents

Kaixin Ma, Di Feng, Alexander Metz +3

The paper introduces MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents that handle multi-image, multi-turn tasks across hundreds of too…

cs.CV2026

SO-Bench: A Structural Output Evaluation of Multimodal LLMs

Di Feng, Kaixin Ma, Feng Nan +9

Multimodal large language models (MLLMs) are increasingly deployed in real-world, agentic settings where outputs must not only be correct, but also conform to predefined data schem…

cs.CL2026

PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice

Yuzhen Shi, Huanghai Liu, Yiran Hu +27

As large language models (LLMs) are increasingly applied to legal domain-specific tasks, evaluating their ability to perform legal work in real-world settings has become essential.…

cs.CV2025

Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents

Zhen Yang, Zi-Yi Dou, Di Feng +13

Developing autonomous agents that effectively interact with Graphic User Interfaces (GUIs) remains a challenging open problem, especially for small on-device models. In this paper,…

cs.CV2025

Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms

Zhangheng Li, Keen You, Haotian Zhang +7

Building a generalist model for user interface (UI) understanding is challenging due to various foundational issues, such as platform diversity, resolution variation, and data limi…