3 papers
cs.CV2026
CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation
Xianjing Han, Yuhan Su, Yang Deng +3
Text-to-video (T2V) generation models have advanced rapidly, yet their ability to represent diverse cultural contexts remains underexplored. Existing benchmarks mainly focus on per…
cs.CV2026
Key-Frame Reasoning with SAM3: Third Place Solution for the MeViS-Text Track of the 8th LSVOS Challenge
Ce Bian, Xusheng He, Jinrong Zhang +3
This report presents a two-stage, training-free solution for the MeViS-Text track of the 8th LSVOS Challenge. The task requires a model to localize and segment the object specified…
cs.CV2026
VOS-Agent: The 1st Place Solution for the 8th LSVOS Challenge (MOSEv2 Track)
Canyang Wu, Jinrong Zhang, Xusheng He +3
Complex video object segmentation requires robust target propagation under severe occlusion, disappearance and reappearance. Although SAM3 provides strong promptable mask propagati…