2 papers
cs.LG2025
Self-Evaluating LLMs for Multi-Step Tasks: Stepwise Confidence Estimation for Failure Detection
Vaibhav Mavi, Shubh Jaroria, Weiqi Sun
Reliability and failure detection of large language models (LLMs) is critical for their deployment in high-stakes, multi-step reasoning tasks. Prior work explores confidence estima…
cs.HC2025
Prompt2Task: Automating UI Tasks on Smartphones from Textual Prompts
Tian Huang, Chun Yu, Weinan Shi +4
UI task automation enables efficient task execution by simulating human interactions with graphical user interfaces (GUIs), without modifying the existing application code. However…