3 papers
cs.CL2026
GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces
Xinyu Geng, Yanjing Xiao, Yuyang Zhang +5
Deep research agents integrate fragmented evidence through multi-step tool use. BrowseComp offers a text-only testbed for such agents, but existing multimodal benchmarks rarely req…
cs.SE2025
TransLibEval: Demystify Large Language Models' Capability in Third-party Library-targeted Code Translation
Pengyu Xue, Kunwu Zheng, Zhen Yang +11
In recent years, Large Language Models (LLMs) have been widely studied in the code translation field on the method, class, and even repository levels. However, most of these benchm…
cs.SE2024
Exploring and Lifting the Robustness of LLM-powered Automated Program Repair with Metamorphic Testing
Pengyu Xue, Linhao Wu, Zhen Yang +8
In recent years, Large language model-powered Automated Program Repair (LAPR) techniques have achieved state-of-the-art bug-fixing performance and have been pervasively applied and…