activity
20232026
most citedKG-Agent: An Efficient Autonomous Agent Framework for Complex Reasoning over Knowledge Graph

13 citations · 53 across the 57 of their papers we have counts for

collaborators
Showing 2026Show all

10 papers · 1 filter

cs.CL2026

ClawGym II: Exploring Black-Box RL on Agent Harness

Huatong Song, Fei Bai, Ming Yang +17

Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment. However, reinforcement learning through compl…

cs.LG2026

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

Fanzhe Meng, Guoxin Chen, Jiale Zhao +6

Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasib…

cs.SE2026

DeNovoSWE: Scaling Long-Horizon Environments for Generating Entire Repositories from Scratch

Jiale Zhao, Guoxin Chen, Fanzhe Meng +4

As the capabilities of LLM-based code agents continue to advance, their expected role is expanding beyond localized bug fixing in existing codebases toward architecting and impleme…

cs.IR2026

Dual-Stream MLP is All You Need for CTR Prediction

Kesha Ou, Zhen Tian, Wayne Xin Zhao +3

Click-through rate (CTR) prediction holds a pivotal role in online advertising and recommendation systems, where even small improvements can significantly boost revenue. Existing r…

cs.SE2026

SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

Huatong Song, Lisheng Huang, Shuang Sun +11

In this technical report, we present SWE-Master, an open-source and fully reproducible post-training framework for building effective software engineering agents. SWE-Master system…

cs.SE2026

SWE-World: Building Software Engineering Agents in Docker-Free Environments

Shuang Sun, Huatong Song, Lisheng Huang +11

Recent advances in large language models (LLMs) have enabled software engineering agents to tackle complex code modification tasks. Most existing approaches rely on execution feedb…