From the 1 of 9 linked papers with an AI index.
9 papers
PDAGENT-BENCH: Characterizing, Grounding, and Architecting LLM/VLM Agents for VLSI Physical Design
Qiufeng Li, Rongqian Chen, Quan Cheng +6
The paper presents PDAGENT-BENCH, a benchmark suite and workflow framework for evaluating large language model and vision‑language model agents on VLSI physical design tasks, cover…
FlowEdit: Information-Theoretic Control of LLM Reasoning Flows for Ill-posed Problems Involving Conflicts
Sizhe Tang, Guangyu Jiang, Yu Li +3
Large Language Models (LLMs) perform strongly on well-specified reasoning tasks with a feasible answer. However, problems encountered in the open world can become ill-posed due to…
IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents
Rongqian Chen, Yu Li, Zeyu Fang +3
Computer-Use Agents (CUAs) leverage large language models to execute GUI operations on desktop environments, yet they generate actions without evaluating action quality, leading to…
Reasoning Knowledge-Gap in Drone Planning via LLM-based Active Elicitation
Zeyu Fang, Beomyeol Yu, Cheng Liu +5
Human-AI joint planning in Unmanned Aerial Vehicles (UAVs) typically relies on control handover when facing environmental uncertainties, which is often inefficient and cognitively…
Knowing When to Ask: Resolving Uncertainty in Human-Robot Joint Planning via Explicit Dialogue and Implicit Intent Cues
Zeyu Fang, Yuxin Lin, Cheng Liu +6
Effective human-robot collaboration in open-world environments requires joint planning under uncertainty about the task, the environment, and the human teammate. Communication is t…
Agent Alpha: Tree Search Unifying Generation, Exploration and Evaluation for Computer-Use Agents
Sizhe Tang, Rongqian Chen, Tian Lan
While scaling test-time compute through trajectory-level sampling has significantly improved Graphical User Interface (GUI) agents, the lack of regressive ability prevents the reus…