Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks
Mengyu Zheng, Kai Han, Boxun Li +13
General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-bench: a generic agent does not b…
cs.LG2026
Mask Is What DLLM Needs: A Masked Data Training Paradigm for Diffusion LLMs
Linrui Ma, Yufei Cui, Kai Han +1
Discrete diffusion models offer global context awareness and flexible parallel generation. However, uniform random noise schedulers in standard DLLM training overlook the highly no…
cs.LG2026
Diffusion In Diffusion: Reclaiming Global Coherence in Semi-Autoregressive Diffusion
Linrui Ma, Yufei Cui, Kai Han +1
One of the most compelling features of global discrete diffusion language models is their global bidirectional contextual capability. However, existing block-based diffusion studie…