collaborators

9 papers

cs.SE2026

A Unified Issue Resolution Benchmark for Requirement Clarification, Planning, and Code Generation for Coding Agents

Xin Zhou, Chun Yong Chong, Kisub Kim +11

Large language model-powered coding agents are increasingly used to modify existing code repositories, for example, by adding features or fixing bugs. Yet existing repository-level…

cs.CV2026

GEB-Bench: Abstract Structures Told in Many Voices

Tong Zhang, Zhiyuan Shi, Yun Peng +1

Can a model look at a river delta and a lightning bolt and see that they share a structure? We introduce GEB-Bench, a benchmark whose unit is an abstract structural motif--self-ref…

cs.CR2026

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs

Tong Zhang, Zexin Li, Simin Chen +1

Jailbreak defenses are essential for protecting large language models (LLMs), but they can also introduce secondary costs that weaken model utility. We present a systematic study o…

cs.LG2026

When Does Recurrence Become an Algorithm? Convergence Selection in Weight-Tied Looped Transformers

Tong Zhang, Junhao Hu, Yun Peng +1

When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm? We answer with four findings from controlled populations on group word pro…

cs.SE2026

What Resolve Rate Hides: Trajectory Structure Diagnostics for Coding Agents

Rui Shu, Chun Yong Chong, Xin Zhou +6

Coding agents are ranked almost entirely by resolve rate: whether their final patch passes the target tests. Yet two agents can reach the same outcome through very different proces…

cs.AI2026

Retrieval-Infused Reasoning Sandbox: A Benchmark for Decoupling Retrieval and Reasoning Capabilities

Shuangshuang Ying, Zheyu Wang, Yunjian Peng +16

Despite strong performance on existing benchmarks, it remains unclear whether large language models can reason over genuinely novel scientific information. Most evaluations score e…