180 citations · 181 across the 13 of their papers we have counts for
6 papers · 1 filter
Conformalized Large Language Models under Configuration Shift
Yuqicheng Zhu, Jialin Yu, Lin Li +7
Conformal prediction (CP) is a distribution-free framework for uncertainty quantification that has recently been adapted to large language models (LLMs), providing prediction sets…
AWM: Answerable Working Memory for Long-Document VQA Agents
Dongzhuoran Zhou, Yuqicheng Zhu, Yule Liu +5
Long-document visual question answering increasingly relies on VLM agents that retrieve candidate pages, inspect page images, write findings to working memory, and synthesize answe…
MV-Bench: Benchmarking Multimodal Large Language Models for Coordinated Multi-View Interface Construction
Yue Zhao, Hongxu Liu, Feiyu Wang +5
Multimodal large language models (MLLMs) are increasingly expected to automate visualization development by generating code directly from visual designs. However, existing evaluati…
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
V Team, Wenyi Hong, Xiaotao Gu +94
We present GLM-5V-Turbo, a step toward native foundation models for multimodal agents. As foundation models are increasingly deployed in real environments, agentic capability depen…
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
Zehai He, Wenyi Hong, Zhen Yang +4
Recent advances in large language models have improved the capabilities of coding agents, yet systematic evaluation of complex, end-to-end website development remains limited. To a…
PlotGen-Bench: Evaluating VLMs on Generating Visualization Code from Diverse Plots across Multiple Libraries
Yi Zhao, Zhen Yang, Shuaiqi Duan +4
Recent advances in vision-language models (VLMs) have expanded their multimodal code generation capabilities, yet their ability to generate executable visualization code from plots…