6 citations · 19 across the 12 of their papers we have counts for
7 papers · 1 filter
DUALVISION: RGB-Infrared Multimodal Large Language Models for Robust Visual Reasoning
Abrar Majeedi, Zhiyuan Ruan, Ziyi Zhao +3
Multimodal large language models (MLLMs) have achieved impressive performance on visual perception and reasoning tasks with RGB imagery, yet they remain fragile under common degrad…
Restore-R1: Efficient Image Restoration Agents via Reinforcement Learning with Multimodal LLM Perceptual Feedback
Jianglin Lu, Yuanwei Wu, Ziyi Zhao +4
Complex image restoration aims to recover high-quality images from inputs affected by multiple degradations such as blur, noise, rain, and compression artifacts. Recent restoration…
SynCDR : Training Cross Domain Retrieval Models with Synthetic Data
Samarth Mishra, Carlos D. Castillo, Hongcheng Wang +2
In cross-domain retrieval, a model is required to identify images from the same semantic category across two visual domains. For instance, given a sketch of an object, a model need…
Lightweight Delivery Detection on Doorbell Cameras
Pirazh Khorramshahi, Zhe Wu, Tianchen Wang +2
Despite recent advances in video-based action recognition and robust spatio-temporal modeling, most of the proposed approaches rely on the abundance of computational resources to a…
VideoSSL: Semi-Supervised Learning for Video Classification
Longlong Jing, Toufiq Parag, Zhe Wu +2
We propose a semi-supervised learning approach for video classification, VideoSSL, using convolutional neural networks (CNN). Like other computer vision tasks, existing supervised…
Layout-induced Video Representation for Recognizing Agent-in-Place Actions
Ruichi Yu, Hongcheng Wang, Ang Li +3
We address the recognition of agent-in-place actions, which are associated with agents who perform them and places where they occur, in the context of outdoor home surveillance. We…