collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward

Chunlin Zhong, Qiuxia Hou, Zhangjun Zhou +5

Video captioning aims to generate comprehensive and coherent descriptions of the video content, contributing to the advancement of both video understanding and generation. However,…

cs.CV2025

Rethinking Detecting Salient and Camouflaged Objects in Unconstrained Scenes

Zhangjun Zhou, Yiping Li, Chunlin Zhong +4

While the human visual system employs distinct mechanisms to perceive salient and camouflaged objects, existing models struggle to disentangle these tasks. Specifically, salient ob…

cs.CV2025

PathVG: A New Benchmark and Dataset for Pathology Visual Grounding

Chunlin Zhong, Shuang Hao, Junhua Wu +5

With the rapid development of computational pathology, many AI-assisted diagnostic tasks have emerged. Cellular nuclei segmentation can segment various types of cells for downstrea…

cs.CV2024

CoLA: Conditional Dropout and Language-driven Robust Dual-modal Salient Object Detection

Shuang Hao, Chunlin Zhong, He Tang

The depth/thermal information is beneficial for detecting salient object with conventional RGB images. However, in dual-modal salient object detection (SOD) model, the robustness a…

cs.CV2024

Tri-modal Confluence with Temporal Dynamics for Scene Graph Generation in Operating Rooms

Diandian Guo, Manxi Lin, Jialun Pei +3

A comprehensive understanding of surgical scenes allows for monitoring of the surgical process, reducing the occurrence of accidents and enhancing efficiency for medical profession…