4 papers
When Semantically Consistent Encoding Meets View-Label Heterogeneity Modeling: A Unified Framework for Incomplete Multi-View Multi-Label Learning
Chengliang Liu, Bo Li, Bob Zhang +3
Incomplete multi-view multi-label learning requires not only robust semantic aggregation from partially observed views, but also label-aware exploitation of view-specific evidence.…
Report of the 8th LSVOS Challenge: Complex and Multimodal Video Object Segmentation
Chang Liu, Henghui Ding, Lingyi Hong +36
This report summarizes the 8th Large-scale Video Object Segmentation (LSVOS) Challenge, held in conjunction with ECCV 2026. The challenge evaluates video segmentation in three comp…
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
Yang-Hao Zhou, Haitian Li, Rexar Lin +12
Recent advances in text-to-audio-video (T2AV) generation have enabled models to synthesize audio-visual videos with multi-participant dialogues. However, existing evaluation benchm…
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
Tian Lan, Yang-Hao Zhou, Zi-Ao Ma +8
Recent advances in deep learning have significantly enhanced generative AI capabilities across text, images, and audio. However, automatically evaluating the quality of these gener…