works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.CV2026

STEP-OPD: Rethinking Output Targets and Internal Dynamics in On-Policy Distillation for Diffusion Models

Qingyan Wei, Guangzhao Li, Xiaobing Tu +5

On-policy distillation (OPD) has become an effective approach for consolidating multiple task-specialized image generation models into a single student. However, existing OPD metho…

cs.CV2026

EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

Zexuan Yan, Yuzhou Wu, Yue Ma +9

The paper introduces EgoGenesis, a simulator that generates controllable egocentric manipulation videos using geometry-aware conditioning mechanisms to augment real robot data and…

cs.CV2026

SpecEdit: Training-Free Acceleration for Diffusion based Image Editing via Semantic Locking

Zhengan Yan, Shikang Zheng, Haoran Qin +9

Diffusion-based image editing offers strong semantic controllability, but remains computationally expensive due to iterative high-resolution denoising over all spatial tokens. Dyna…

cs.AI2025

AgentBay: A Hybrid Interaction Sandbox for Seamless Human-AI Intervention in Agentic Systems

Yun Piao, Hongbo Min, Hang Su +28

The rapid advancement of Large Language Models (LLMs) is catalyzing a shift towards autonomous AI Agents capable of executing complex, multi-step tasks. However, these agents remai…

cs.CV2025

GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents

Xianhang Ye, Yiqing Li, Wei Dai +8

Existing GUI grounding methods often struggle with fine-grained localization in high-resolution screenshots. To address this, we propose GUI-ARP, a novel framework that enables ada…

cs.LG2025

TGPO: Tree-Guided Preference Optimization for Robust Web Agent Reinforcement Learning

Ziyuan Chen, Zhenghui Zhao, Zhangye Han +7

With the rapid advancement of large language models and vision-language models, employing large models as Web Agents has become essential for automated web interaction. However, tr…