6 citations · 8 across the 4 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024★ 6 cited
MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Kaining Ying, Fanqing Meng, Jin Wang +19
Large Vision-Language Models (LVLMs) show significant strides in general-purpose multimodal applications such as visual dialogue and embodied navigation. However, existing multimod…
cs.CV2023
Human-to-Human Interaction Detection
Zhenhua Wang, Kaining Ying, Jiajun Meng +1
A comprehensive understanding of interested human-to-human interactions in video streams, such as queuing, handshaking, fighting and chasing, is of immense importance to the survei…
cs.CV2023★ 1 cited
CTVIS: Consistent Training for Online Video Instance Segmentation
Kaining Ying, Qing Zhong, Weian Mao +7
The discrimination of instance embeddings plays a vital role in associating instances across time for online video instance segmentation (VIS). Instance embedding learning is direc…