3 papers
cs.SE2026
SWE-Explore: Benchmarking How Coding Agents Explore Repositories
Shaoqiu Zhang, Yuhang Wang, Jialiang Liang +8
Repository-level coding benchmarks such as SWE-bench have driven a rapid surge in the capabilities of coding agents. Yet they usually treat coding tasks as a holistic, binary predi…
cs.CL2026
Semantic Bridging Domains: Pseudo-Source as Test-Time Connector
Xizhong Yang, Huiming Wang, Ning Xu +1
Distribution shifts between training and testing data are a critical bottleneck limiting the practical utility of models, especially in real-world test-time scenarios. To adapt mod…
cs.LG2025
MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs
Chenchen Zhao, Zhengyuan Shi, Xiangyu Wen +19
The emergence of multimodal large language models (MLLMs) presents promising opportunities for automation and enhancement in Electronic Design Automation (EDA). However, comprehens…