5 papers
FabriVLA: A Lightweight Vision-Language-Action Model with Conformal Action Chunk Uncertainty
Shiyuan Yang, Borong Zhang, Jizheng Zhang +5
Vision-Language-Action (VLA) models have become a leading paradigm for general purpose robotic manipulation, but their computational cost and limited uncertainty awareness hinder p…
BoxCtrl: 3D-Aware Visual Prompting for Geometric Image Editing
Feifei Wang, Shiyuan Yang, Xiaoyu Li +1
As instruction-based editing models and multimodal large language models advance, diverse image editing tasks have become feasible. However, achieving precise and consistent geomet…
OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control
Yukun Wang, Ruihuang Li, Jiale Tao +7
Video fundamentally intertwines two crucial axes: the dynamic content of a scene and the camera motion through which it is observed. However, existing generation models often entan…
EffectMaker: Unifying Reasoning and Generation for Customized Visual Effect Creation
Shiyuan Yang, Ruihuang Li, Jiale Tao +3
Visual effects (VFX) are essential for enhancing the expressiveness and creativity of video content, yet producing high-quality effects typically requires expert knowledge and cost…
MTV-Inpaint: Multi-Task Long Video Inpainting
Shiyuan Yang, Zheng Gu, Liang Hou +4
Video inpainting involves modifying local regions within a video, ensuring spatial and temporal consistency. Most existing methods focus primarily on scene completion (i.e., fillin…