18 papers
Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models
Jiazheng Xing, Hangjie Yuan, Lingling Cai +9
Connector-based video unified models have demonstrated strong capability in instruction-grounded video synthesis, but integrating a large high-fidelity generator into the unified t…
ActionMap: Robot Policy Learning via Voxel Action Heatmap
Pei Yang, Hai Ci, Yanzhe Chen +3
Vision-language-action (VLA) models have advanced rapidly across backbones, training recipes, and data scale, yet the action decoder, which converts the backbone's hidden state int…
World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy
Xiaokang Liu, Zechen Bai, Hai Ci +2
Reinforcement learning (RL) can refine Vision-Language-Action (VLA) policies beyond behavior cloning, but real-world RL remains expensive due to extensive rollouts, resets, supervi…
UENR-600K: A Large-Scale Physically Grounded Dataset for Nighttime Video Deraining
Pei Yang, Hai Ci, Beibei Lin +2
Nighttime video deraining is uniquely challenging because raindrops interact with artificial lighting. Unlike daytime white rain, nighttime rain takes on various colors and appears…
RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video
Haiyang Mei, Qiming Huang, Hai Ci +1
Accurate robot segmentation is a fundamental capability for robotic perception. It enables precise visual servoing for VLA systems, scalable robot-centric data augmentation, accura…
LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation
Jiazheng Xing, Fei Du, Hangjie Yuan +7
Recent advances in diffusion models have significantly improved text-to-video generation, enabling personalized content creation with fine-grained control over both foreground and…