2 citations · 2 across the 5 of their papers we have counts for
Showing 2026 · cs.CVShow all
3 papers · 2 filters
cs.CV2026
Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models
Yuancheng Wei, Linli Yao, Lei Li +4
Multimodal reward models have advanced substantially in text and image domains, yet progress in video understanding reward modeling remains severely limited by the lack of robust e…
cs.CV2026
DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning
Yuancheng Wei, Haojie Zhang, Linli Yao +7
Image Difference Captioning (IDC) generates natural language descriptions that precisely identify differences between two images, serving as a key benchmark for fine-grained change…
cs.CV2026
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
Linli Yao, Yuancheng Wei, Yaojie Zhang +12
This paper proposes Omni Dense Captioning, a novel task designed to generate continuous, fine-grained, and structured audio-visual narratives with explicit timestamps. To ensure de…