collaborators

24 papers

cs.CV2026

Is It Time for the Renaissance of Salient Object Detection in the Era of MLLMs?

Wenzhuo Zhao, Xiuzhi Li, Zhongkuan Mao +6

The zero-shot capabilities of multimodal large language models (MLLMs) are pushing salient object detection (SOD) beyond task-specific supervision. To disentangle MLLMs beyond conv…

cs.CV2026

RDVSv2: A Large-scale Benchmark for RGB-D Video Salient Object Detection

Tianyu Li, Jiahao He, Keren Fu +1

We introduce RDVSv2, a large-scale benchmark for RGB-D video salient object detection (RGB-D VSOD) with dense frame-level annotations. Existing datasets in this emerging field are…

cs.CV2026

CamoSAM2: SAM2-oriented Prompt Auto-Refinement for Video Camouflaged Object Detection

Xin Zhang, Keren Fu, Qijun Zhao

The Segment Anything Model 2 (SAM2), a prompt-guided video foundation model, has remarkably performed in video object segmentation, drawing significant attention in the community.…

cs.CV2026

Attend to Anything: Foundation Model for Unified Human Attention Modeling

Wenzhuo Zhao, Ronghao Xian, Keren Fu +1

Existing human attention (saliency) modeling methods persist as highly fragmented across modalities, scenes, and task formulations. Consequently, even with increasing model capacit…

cs.LG2026

Physics-Informed Generative Solver: Bridging Data-Driven Priors and Conservation Laws for Stable Spatiotemporal Field Reconstruction

Ziyuan Zhu, Keyu Hu, Zhifei Chen +10

Reconstructing continuous physical fields from sparse measurements is a central inverse problem, but data-driven generative models can produce states that violate governing dynamic…

cs.CV2026

DAPL: Integration of Positive and Negative Descriptions in Text-Based Person Search

Yuchuan Deng, Zhanpeng Hu, Zijie Xin +2

Text-based person search (TBPS) aims to retrieve specific images of individuals from large datasets using textual descriptions. Existing TBPS methods focus primarily on identifying…