9 papers · 1 filter
FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Search
James Xu Zhao, Hui Chen, Bryan Hooi +1
Agentic search requires language model agents to explore many sources and answer complex information-seeking questions. Scaling test-time compute is a promising way to improve thes…
ConvApparel: A Benchmark Dataset and Validation Framework for User Simulators in Conversational Recommenders
Ofer Meshi, Krisztian Balog, Sally Goldman +5
The promise of LLM-based user simulators to improve conversational AI is hindered by a critical "realism gap," leading to systems that are optimized for simulated interactions, but…
Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning
Tianyi Wu, Jingwei Ni, Bryan Hooi +5
Instruction fine-tuning (IFT) can increase the informativeness of large language models (LLMs), but may reduce their truthfulness. This trade-off arises because IFT steers LLMs to…
How Does Response Length Affect Long-Form Factuality
James Xu Zhao, Jimmy Z. J. Liu, Bryan Hooi +1
Large language models (LLMs) are widely used for long-form text generation. However, factual errors in the responses would undermine their reliability. Despite growing attention to…
Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models
Zhiyuan Hu, Yibo Wang, Hanze Dong +5
Large reasoning models (LRMs) already possess a latent capacity for long chain-of-thought reasoning. Prior work has shown that outcome-based reinforcement learning (RL) can inciden…
Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation
Zhiyuan Hu, Shiyun Xiong, Yifan Zhang +5
Recent advancements in visual language models (VLMs) have notably enhanced their capabilities in handling complex Graphical User Interface (GUI) interaction tasks. Despite these im…