10 papers
FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation
Hao Liu, Chenghuan Huang, Ye Huang +6
Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention r…
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation
Peiwen Zhang, Yufan Deng, Shangkun Sun +11
Video generation models have emerged as a promising paradigm for embodied world simulation. However, both general-domain video generators and robot-specific data fine-tuned models…
HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining
Juncheng Ma, Jianxin Bi, Yufan Deng +19
Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tighter data bottleneck. Teleoperated real-robot trajectories remai…
Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time
Ruizhi Zhang, Ye Huang, Yuangang Pan +6
While artificial intelligence has mastered structured games like chess and Go, vision-language agents still struggle in visually-driven 3D games without access to game states. Exis…
Coding with Eyes: Visual Feedback Unlocks Reliable GUI Code Generating and Debugging
Zhilin Liu, Ye Huang, Ting Xie +3
Recent advances in Large Language Model (LLM)-based agents have shown remarkable progress in code generation. However, current agent methods mainly rely on text-output-based feedba…
Tuning-Free Adaptive Style Incorporation for Structure-Consistent Text-Driven Style Transfer
Yanqi Ge, Jiaqi Liu, Qingnan Fan +6
In this work, we target the task of text-driven style transfer in the context of text-to-image (T2I) diffusion models. The main challenge is consistent structure preservation while…