3 papers
cs.CV2026
Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos
Runze Xu, Yiluo Zhang, Jian Wang +2
Training generalist Vision-Language-Action(VLA) models typically requires massive, diverse robotic datasets with high-fidelity action annotations. While egocentric human manipulati…
cs.CV2026
Video2Code: Generating Interactive Webpages from UI Videos via Action-Aware Revisit
Mingde Xu, Zhen Yang, Yan Wang +7
UI videos provide a natural input for generating interactive webpages, as they capture both webpage appearance and action-triggered state transitions. However, directly applying vi…
cs.CV2025
Towards Better Optimization For Listwise Preference in Diffusion Models
Jiamu Bai, Xin Yu, Meilong Xu +6
Reinforcement learning from human feedback (RLHF) has proven effectiveness for aligning text-to-image (T2I) diffusion models with human preferences. Although Direct Preference Opti…