5 papers
Supervise What Survives: Geometry-Guided VLA Adaptation from Synthetic Robot Videos
Danze Chen, Yanzhe Chen, Qiming Huang +3
Vision-Language-Action (VLA) models require large-scale video-action pairs, yet real teleoperation remains scarce. While generated robot videos offer a scalable alternative, existi…
RobotSeg: A Model and Dataset for Segmenting Robots in Image and Video
Haiyang Mei, Qiming Huang, Hai Ci +1
Accurate robot segmentation is a fundamental capability for robotic perception. It enables precise visual servoing for VLA systems, scalable robot-centric data augmentation, accura…
Structure-Aware Feature Rectification with Region Adjacency Graphs for Training-Free Open-Vocabulary Semantic Segmentation
Qiming Huang, Hao Ai, Jianbo Jiao
Benefiting from the inductive biases learned from large-scale datasets, open-vocabulary semantic segmentation (OVSS) leverages the power of vision-language models, such as CLIP, to…
Audio-Visual Separation with Hierarchical Fusion and Representation Alignment
Han Hu, Dongheng Lin, Qiming Huang +3
Self-supervised audio-visual source separation leverages natural correlations between audio and vision modalities to separate mixed audio signals. In this work, we first systematic…
What Can We Learn from Harry Potter? An Exploratory Study of Visual Representation Learning from Atypical Videos
Qiyue Sun, Qiming Huang, Yang Yang +2
Humans usually show exceptional generalisation and discovery ability in the open world, when being shown uncommon new concepts. Whereas most existing studies in the literature focu…