3 papers
cs.AI2026
DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis
Qiaohong Zhang, Weihao Ye, Jialong Chen +7
Autonomous data analysis agents are increasingly expected to conduct exploratory analysis with limited human guidance about data. However, existing benchmarks typically evaluate su…
cs.SE2026
SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration
Jialong Chen, Xander Xu, Hu Wei +2
Large language model (LLM)-powered agents have demonstrated strong capabilities in automating software engineering tasks such as static bug fixing. However, in the real world, the…
cs.LG2025
THESAURUS: Contrastive Graph Clustering by Swapping Fused Gromov-Wasserstein Couplings
Bowen Deng, Tong Wang, Lele Fu +3
Graph node clustering is a fundamental unsupervised task. Existing methods typically train an encoder through selfsupervised learning and then apply K-means to the encoder output.…