3 papers
cs.CV2026
Report of the 8th LSVOS Challenge: Complex and Multimodal Video Object Segmentation
Chang Liu, Henghui Ding, Lingyi Hong +36
This report summarizes the 8th Large-scale Video Object Segmentation (LSVOS) Challenge, held in conjunction with ECCV 2026. The challenge evaluates video segmentation in three comp…
cs.CV2026
MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge
Liangtao Shi, Jinxia Xie, Xiantao Hu +1
In this technical report, we present a training-free framework for audio-guided video object segmentation, which integrates Multimodal Large Language Models (MLLMs) with SAM-based…
cs.CV2024
Robust Tracking via Mamba-based Context-aware Token Learning
Jinxia Xie, Bineng Zhong, Qihua Liang +3
How to make a good trade-off between performance and computational cost is crucial for a tracker. However, current famous methods typically focus on complicated and time-consuming…