8 papers
SCP-NL2TL: Selective Conformal Prediction with Semantic Verification for Natural Language to Temporal Logic Specifications
Yixuan Wang, Licheng Luo, Yu Fu +3
Translating natural language instructions into machine-interpretable formal specifications enables robots and autonomous systems to plan, reason, and formally verify their behavior…
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling
Peng Kuang, Haibo Jin, Xiaoyu Han +5
Process Reward Models (PRMs) have been proven to be highly effective in guiding test-time scaling (TTS) methods, which significantly boost the capabilities of LLM-based multi-agent…
SeClaw: Spec-Driven Security Task Synthesis for Evaluating Autonomous Agents
Hao Cheng, Changtao Miao, Tianle Song +21
Autonomous LLM agents increasingly operate in stateful environments where they access tools, files, memory, and external services. While such capabilities enable complex real-world…
PaintBench: Deterministic Evaluation of Precise Visual Editing
Kai Xu, Ellis Brown, Shrikar Madhu +3
While current multimodal models are proficient at open-ended visual editing, executing precise single-answer edits remains an important obstacle. To probe this challenge, we introd…
Optimal Aggregation of LLM and PRM Signals for Efficient Test-Time Scaling
Peng Kuang, Yanli Wang, Xiaoyu Han +3
Process reward models (PRMs) are a cornerstone of test-time scaling (TTS), designed to verify and select the best responses from large language models (LLMs). However, this promise…
IUQ: Interrogative Uncertainty Quantification for Long-Form Large Language Model Generation
Haozhi Fan, Jinhao Duan, Kaidi Xu
Despite the rapid advancement of Large Language Models (LLMs), uncertainty quantification in LLM generation is a persistent challenge. Although recent approaches have achieved stro…