From the 1 of 6 linked papers with an AI index.
6 papers
STEP-OPD: Rethinking Output Targets and Internal Dynamics in On-Policy Distillation for Diffusion Models
Qingyan Wei, Guangzhao Li, Xiaobing Tu +5
On-policy distillation (OPD) has become an effective approach for consolidating multiple task-specialized image generation models into a single student. However, existing OPD metho…
EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE
Zexuan Yan, Yuzhou Wu, Yue Ma +9
The paper introduces EgoGenesis, a simulator that generates controllable egocentric manipulation videos using geometry-aware conditioning mechanisms to augment real robot data and…
SpecEdit: Training-Free Acceleration for Diffusion based Image Editing via Semantic Locking
Zhengan Yan, Shikang Zheng, Haoran Qin +9
Diffusion-based image editing offers strong semantic controllability, but remains computationally expensive due to iterative high-resolution denoising over all spatial tokens. Dyna…
AgentBay: A Hybrid Interaction Sandbox for Seamless Human-AI Intervention in Agentic Systems
Yun Piao, Hongbo Min, Hang Su +28
The rapid advancement of Large Language Models (LLMs) is catalyzing a shift towards autonomous AI Agents capable of executing complex, multi-step tasks. However, these agents remai…
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
Xianhang Ye, Yiqing Li, Wei Dai +8
Existing GUI grounding methods often struggle with fine-grained localization in high-resolution screenshots. To address this, we propose GUI-ARP, a novel framework that enables ada…
TGPO: Tree-Guided Preference Optimization for Robust Web Agent Reinforcement Learning
Ziyuan Chen, Zhenghui Zhao, Zhangye Han +7
With the rapid advancement of large language models and vision-language models, employing large models as Web Agents has become essential for automated web interaction. However, tr…