3 papers
cs.CV2025
RefOnce: Distilling References into a Prototype Memory for Referring Camouflaged Object Detection
Yu-Huan Wu, Zi-Xuan Zhu, Yan Wang +2
Referring Camouflaged Object Detection (Ref-COD) segments specified camouflaged objects in a scene by leveraging a small set of referring images. Though effective, current systems…
cs.CV2025
FOCUS: Efficient Keyframe Selection for Long Video Understanding
Zirui Zhu, Hailun Xu, Yang Luo +4
Multimodal large language models (MLLMs) represent images and video frames as visual tokens. Scaling from single images to hour-long videos, however, inflates the token budget far…
cs.CV2025
GAPNet: A Lightweight Framework for Image and Video Salient Object Detection via Granularity-Aware Paradigm
Yu-Huan Wu, Wei Liu, Zi-Xuan Zhu +3
Recent salient object detection (SOD) models predominantly rely on heavyweight backbones, incurring substantial computational cost and hindering their practical application in vari…