8 papers
An Evolutionary Agentic Approach for Open-ended Image Quality Perception
Zhenchen Tang, Bo Peng, Zichuan Wang +4
Generative models are rapidly expanding image quality assessment (IQA) beyond traditional fidelity factors to emerging dimensions such as physical plausibility and text-rendering c…
Transsion's Speaker-Attributed Multilingual ASR System for the MLC-SLM 2026 Challenge
Zhecheng Ren, Xuanji He, Xiaoxiao Li +5
This paper presents the Transsion Speech Team submission to Task 1 of the MLC-SLM 2026 Challenge, which focuses on speaker-attributed transcription for multilingual conversational…
The tttAI System for the TSA-ASR Task of the SmartGlasses Challenge 2026
Xuanji He, Gaoyang Dong, Xiaoxiao Li +2
This paper presents the tttAI system submitted to the TSA-ASR task of the SmartGlasses Challenge 2026, evaluated on both two-person dialogues (Track 1) and multi-party meetings (Tr…
A Trajectory-Driven Spatio-Temporal Refinement Solution for CVPR 2026 8th UG2+ Challenge Track 3: DOST
Hongzhen Li, Miao Yu, Leilei Cao +3
In this work, we present our solution for the 8th UG2+ Challenge (CVPR 2026) Track 3: Dynamic Object Segmentation in Turbulence (DOST). Our method is built upon the strong baseline…
LSVOS 2025 Challenge Report: Recent Advances in Complex Video Object Segmentation
Chang Liu, Henghui Ding, Kaining Ying +46
This report presents an overview of the 7th Large-scale Video Object Segmentation (LSVOS) Challenge held in conjunction with ICCV 2025. Besides the two traditional tracks of LSVOS…
Enhancing Sa2VA for Referent Video Object Segmentation: 2nd Solution for 7th LSVOS RVOS Track
Ran Hong, Feng Lu, Leilei Cao +3
Referential Video Object Segmentation (RVOS) aims to segment all objects in a video that match a given natural language description, bridging the gap between vision and language un…