3 papers
cs.CL2026
FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Search
James Xu Zhao, Hui Chen, Bryan Hooi +1
Agentic search requires language model agents to explore many sources and answer complex information-seeking questions. Scaling test-time compute is a promising way to improve thes…
cs.AI2026
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
James Xu Zhao, Bryan Hooi, See-Kiong Ng
Test-time scaling increases inference-time computation through longer reasoning chains and has shown strong performance gains across many domains. However, frontier models still su…
cs.CL2025
How Does Response Length Affect Long-Form Factuality
James Xu Zhao, Jimmy Z. J. Liu, Bryan Hooi +1
Large language models (LLMs) are widely used for long-form text generation. However, factual errors in the responses would undermine their reliability. Despite growing attention to…