14 papers
Competitive Memory Readout for Robust Video Object Segmentation: 2nd Place Technical Report for the MOSEv2 Track of the 8th LSVOS Challenge
Mingqi Gao, Sijie Li, Jungong Han
We present our solution for the MOSEv2 track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge at ECCV 2026. The challenge evaluates robust video object segmentati…
Hy-Embodied-VLM-1.0: Efficient Physical-World Agents
Ziyi Wang, Xumin Yu, Yongming Rao +19
Building capable embodied agents requires not only multimodal perception and understanding, but also agentic capabilities for reasoning about actions, adapting to evolving situatio…
Token Predictors Are Not Planners: Building Physically Grounded Causal Reasoners
Zheng Lu, Mingqi Gao, Qinlei Xie +8
Current benchmarks for embodied vision-language planning often favor linguistic next-token prediction over physically grounded next-state reasoning. This rewards models that mimic…
Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding
Chang Liu, Henghui Ding, Nikhila Ravi +40
This report summarizes the objectives, datasets, and top-performing methodologies of the 2026 Pixel-level Video Understanding in the Wild (PVUW) Challenge, hosted at CVPR 2026, whi…
Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment
Jingkun Chen, Ruoshi Xu, Mingqi Gao +2
Point-Vision-Language Models promise to empower embodied agents with executable spatial reasoning, yet they frequently succumb to geometric hallucination where predicted 3D structu…
UniSurgSAM: A Unified Promptable Model for Reliable Surgical Video Segmentation
Haofeng Liu, Ziyue Wang, Alex Y. W. Kong +6
Surgical video segmentation is fundamental to computer-assisted surgery. In practice, surgeons need to dynamically specify targets throughout extended procedures, using heterogeneo…