9 papers · 1 filter
Auto-Configuring Scientific Simulators with Lightweight Coding-Agent Adapters
Matthew Ho, Brian Liu, Jixuan Chen +2
Configuring an advanced scientific simulator, translating a modeling goal into a valid, runnable input deck, is a persistent bottleneck that costs domain scientists hours to days.…
OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces
Xiaozhe Li, Jixuan Chen, Xinyu Fang +4
Large Language Models (LLMs) have demonstrated remarkable capabilities in reasoning and tool use. However, the fundamental cognitive faculties essential for problem solving, includ…
DeliveryBench: Can Agents Earn Profit in Real World?
Lingjun Mao, Jiawei Ren, Kun Zhou +3
LLMs and VLMs are increasingly deployed as embodied agents, yet existing benchmarks largely revolve around simple short-term tasks and struggle to capture rich realistic constraint…
COMMA: A Communicative Multimodal Multi-Agent Benchmark
Timothy Ossowski, Danyal Maqbool, Jixuan Chen +3
The rapid advances of multimodal agents built on large foundation models have largely overlooked their potential for language-based communication between agents in collaborative ta…
Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
Tianbao Xie, Jiaqi Deng, Xiaochuan Li +12
Graphical user interface (GUI) grounding, the ability to map natural language instructions to specific actions on graphical user interfaces, remains a critical bottleneck in comput…
OpenCUA: Open Foundations for Computer-Use Agents
Xinyuan Wang, Bowen Wang, Dunjie Lu +39
Vision-language models have demonstrated impressive capabilities as computer-use agents (CUAs) capable of automating diverse computer tasks. As their commercial potential grows, cr…