24 citations · 24 across the 3 of their papers we have counts for
6 papers · 1 filter
Self-Evolving Agentic Image Restoration via Deliberate Planning and Intuitive Execution
Shuang Cui, Fan Ji, Guanglong Sun +4
Real-world image restoration (IR) remains challenging due to complex and coupled degradations. While recent agentic IR frameworks leverage Large Language Models for flexible tool p…
Towards Artwork Explanation in Large-scale Vision Language Models
Kazuki Hayashi, Yusuke Sakai, Hidetaka Kamigaito +2
Large-scale Vision-Language Models (LVLMs) output text from images and instructions, demonstrating capabilities in text generation and comprehension. However, it has not been clari…
Synergistic Prompting for Robust Visual Recognition with Missing Modalities
Zhihui Zhang, Luanyuan Dai, Qika Lin +5
Large-scale multi-modal models have demonstrated remarkable performance across various visual recognition tasks by leveraging extensive paired multi-modal training data. However, i…
Toward Realistic Camouflaged Object Detection: Benchmarks and Method
Zhimeng Xin, Tianxu Wu, Shiming Chen +5
Camouflaged object detection (COD) primarily relies on semantic or instance segmentation methods. While these methods have made significant advancements in identifying the contours…
DiffX: Guide Your Layout to Cross-Modal Generative Modeling
Zeyu Wang, Jingyu Lin, Yifei Qian +8
Diffusion models have made significant strides in language-driven and layout-driven image generation. However, most diffusion models are limited to visible RGB image generation. In…
VISTA: A Visual and Textual Attention Dataset for Interpreting Multimodal Models
Harshit, Tolga Tasdizen
The recent developments in deep learning led to the integration of natural language processing (NLP) with computer vision, resulting in powerful integrated Vision and Language Mode…