collaborators

7 papers

cs.CL2026

GISA: A Benchmark for General Information-Seeking Assistant

Yutao Zhu, Xingshuo Zhang, Maosen Zhang +9

The advancement of large language models (LLMs) has significantly accelerated the development of search agents capable of autonomously gathering information through multi-turn web…

cs.LG2026

ContextBench: A Benchmark for Context Retrieval in Coding Agents

Han Li, Letian Zhu, Bohan Zhang +7

LLM-based coding agents have shown strong performance on automated issue resolution benchmarks, yet existing evaluations largely focus on final task success, providing limited insi…

cs.SE2026

Prometheus: Towards Long-Horizon Codebase Navigation for Repository-Level Problem Solving

Yue Pan, Zimin Chen, Siyu Lu +8

Large Language Models (LLMs) have shown remarkable capabilities in automating software engineering tasks, spurring the emergence of coding agents that scaffold LLMs with external t…

cs.AI2026

SimGym: Traffic-Grounded Browser Agents for Offline A/B Testing in E-Commerce

Alberto Castelo, Zahra Zanjani Foumani, Ailin Fan +17

A/B testing remains the gold standard for evaluating e-commerce UI changes, yet it diverts traffic, takes weeks to reach significance, and risks harming user experience. We introdu…

cs.SE2025

ReVeal: Self-Evolving Code Agents via Reliable Self-Verification

Yiyang Jin, Kunzhao Xu, Hang Li +4

Reinforcement learning with verifiable rewards (RLVR) has advanced the reasoning capabilities of large language models. However, existing methods rely solely on outcome rewards, wi…

cs.DB2025

RubikSQL: Lifelong Learning Agentic Knowledge Base as an Industrial NL2SQL System

Zui Chen, Han Li, Xinhao Zhang +12

We present RubikSQL, a novel NL2SQL system designed to address key challenges in real-world enterprise-level NL2SQL, such as implicit intents and domain-specific terminology. Rubik…