3 papers
cs.IR2026
Projecting BrowseComp-Plus onto ClimbMix: Toward More Realistic Corpora for Agentic Search
Sahel Sharifymoghaddam, Lingwei Gu, Yijun Ge +1
The BrowseComp-Plus benchmark disentangled the evaluation of agentic search by replacing opaque web search with a fixed corpus, so that an agent's role can be separated from the re…
cs.CL2026
NanoKnow: How to Know What Your Language Model Knows
Lingwei Gu, Nour Jedidi, Jimmy Lin
How do large language models (LLMs) know what they know? Answering this question has been difficult because pre-training data is often a "black box" - unknown or inaccessible. The…
cs.CV2024
NoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries
Tao Wu, Chuhao Zhou, Yen Heng Wong +2
The rapid advancement of Vision-Language Models (VLMs) has significantly advanced the development of Embodied Question Answering (EQA), enhancing agents' abilities in language unde…