collaborators

9 papers

cs.CR2026

CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild

Jian Yang, Haau-Sing Li, Shawn Guo +10

As large language models (LLMs) continue to advance in coding capabilities, their potential in cybersecurity has drawn increasing research attention, with closed-source LLMs (e.g.,…

cs.CL2026

WebWorld: The Browser as a World Model for Self-Improving Web Code

Jiajun Wu, Jian Yang, Yaxin Du +7

VLM-driven self-improvement of web code has a structural flaw: the model that proposes the repair is the model that judges it, and visual plausibility under that judge is a poor pr…

cs.AI2026

MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection

Haowen Wang, Yaxin Du, Jian Yang +9

Mid-training has become an important stage in modern LLM development, using large-scale curated mixtures to strengthen capabilities before final post-training. Its data selection p…

cs.AI2026

AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security

Dongrui Liu, Yu Li, Zhonghao Yang +47

Modern open-world agents such as OpenClaw exhibit powerful cross-environment execution capabilities yet introduce broad new safety risk sources. Meanwhile, advanced frontier AI mod…

cs.SE2026

HTMLCure: Turning Browser Experience into State Guided Repair for Interactive HTML

Jiajun Wu, Jian Yang, Tuney Zheng +4

LLMs can now produce full HTML pages, but many of those pages are only superficially correct: they render once, then fail under scroll, hover, click, resize, or gameplay. Evaluatio…

cs.CV2026

SGMD: Score Gradient Matching Distillation for Few-Step Video Diffusion Distillation

Zhuguanyu Wu, Ruihao Gong, Yang Yong +5

Distribution Matching Distillation (DMD) is a widely used paradigm for accelerating inference in few-step video diffusion models. However, DMD-style video distillation faces two co…