11 papers · 1 filter
Report of the 8th LSVOS Challenge: Complex and Multimodal Video Object Segmentation
Chang Liu, Henghui Ding, Lingyi Hong +36
This report summarizes the 8th Large-scale Video Object Segmentation (LSVOS) Challenge, held in conjunction with ECCV 2026. The challenge evaluates video segmentation in three comp…
DeSeG: Decoupling Semantic Intent and Geometric Constraints for Physically Plausible Human-Scene Interaction
Jiakun Li, Zhe Li, Wenqiang Wu +4
Synthesizing physically plausible human-scene interactions (HSI) remains a critical challenge in computer vision and the development of human avatars. Although recent generative mo…
Learning Trajectory-Aware Multimodal Large Language Models for Video Reasoning Segmentation
Jingnan Luo, Mingqi Gao, Jun Liu +2
The prosperity of Multimodal Large Language Models (MLLMs) has stimulated the demand for video reasoning segmentation, which aims to segment video objects based on human instructio…
ArtiWorld: LLM-Driven Articulation of 3D Objects in Scenes
Yixuan Yang, Luyang Xie, Zhen Luo +4
Building interactive simulators and scalable robot-learning environments requires a large number of articulated assets. However, most existing 3D assets in simulation are rigid, an…
Leveraging Geometric Priors for Unaligned Scene Change Detection
Ziling Liu, Ziwei Chen, Mingqi Gao +2
Unaligned Scene Change Detection aims to detect scene changes between image pairs captured at different times without assuming viewpoint alignment. To handle viewpoint variations,…
SCORP: Scene-Consistent Object Refinement via Proxy Generation and Tuning
Ziwei Chen, Ziling Liu, Zitong Huang +2
Viewpoint missing of objects is common in scene reconstruction, as camera paths typically prioritize capturing the overall scene structure rather than individual objects. This make…