From the 1 of 4 linked papers with an AI index.
4 papers
How Benchmarks Mis-Score Computer-Use Agents
Zihan Dong, Zhiyuan Ma, Zekun Wang +5
The paper examines how current benchmarks for computer-use agents often give inaccurate scores due to issues in task design, trajectory observation, scoring, and reporting, and pro…
ManuRAG: Multi-modal Retrieval Augmented Generation for Manufacturing Question Answering (Early Version)
Yunqing Li, Zihan Dong, Farhad Ameri +1
The evolution of digital manufacturing requires intelligent Question Answering (QA) systems that can seamlessly integrate and analyze complex multi-modal data, such as text, images…
Towards General Computer Control with Hierarchical Agents and Multi-Level Action Spaces
Zihan Dong, Xinyu Fan, Zixiang Tang +1
Controlling desktop applications via software remains a fundamental yet under-served problem. Existing multi-modal large language models (MLLMs) ingest screenshots and task instruc…
Does Learning Mathematical Problem-Solving Generalize to Broader Reasoning?
Ruochen Zhou, Minrui Xu, Shiqi Chen +5
There has been a growing interest in enhancing the mathematical problem-solving (MPS) capabilities of large language models. While the majority of research efforts concentrate on c…