5 papers · 1 filter
SAFER: Risk-Constrained Sample-then-Filter in Large Language Models
Qingni Wang, Yue Fan, Xin Eric Wang
As large language models (LLMs) are increasingly deployed in risk-sensitive applications such as real-world open-ended question answering (QA), ensuring the trustworthiness of thei…
SafeGround: Know When to Trust GUI Grounding Models via Uncertainty Calibration
Qingni Wang, Yue Fan, Xin Eric Wang
Graphical User Interface (GUI) grounding aims to translate natural language instructions into executable screen coordinates, enabling automated GUI interaction. Nevertheless, incor…
Cross-Modal Memory Compression for Efficient Multi-Agent Debate
Jing Wu, Yue Sun, Tianpei Xie +7
Multi-agent debate can improve reasoning quality and reduce hallucinations, but it incurs rapidly growing context as debate rounds and agent count increase. Retaining full textual…
PerPilot: Personalizing VLM-based Mobile Agents via Memory and Exploration
Xin Wang, Zhiyao Cui, Hao Li +10
Vision language model (VLM)-based mobile agents show great potential for assisting users in performing instruction-driven tasks. However, these agents typically struggle with perso…
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
Kinjal Basu, Ibrahim Abdelaziz, Kiran Kate +10
The resurgence of autonomous agents built using large language models (LLMs) to solve complex real-world tasks has brought increased focus on LLMs' fundamental ability of tool or f…