activity
20242026
collaborators

9 papers

cs.CV2026

Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding

Hang Wu, Sherin Mary Mathews, Yujun Cai +2

Online streaming video understanding requires models to process continuous visual inputs and respond to user queries in real time, where the unbounded stream and unpredictable quer…

cs.CV2026

Efficient Degradation-agnostic Image Restoration via Channel-Wise Functional Decomposition and Manifold Regularization

Bin Ren, Yawei Li, Xu Zheng +6

Degradation-agnostic image restoration aims to handle diverse corruptions with one unified model, but faces fundamental challenges in balancing efficiency and performance across di…

cs.CV2026

Any Image Restoration via Efficient Spatial-Frequency Degradation Adaptation

Bin Ren, Eduard Zamfir, Zongwei Wu +7

Restoring multiple degradations efficiently via just one model has become increasingly significant and impactful, especially with the proliferation of mobile devices. Traditional s…

cs.CV2025

Parameter-Efficient CLIP Adaptation for 3D Understanding via Unified Tokenization

Guofeng Mei, Bin Ren, Qinfeng Xiao +8

Vision-language models, such as CLIP, encode rich semantic knowledge through large-scale image-text pretraining. Reusing these models for 3D understanding is highly desirable, beca…

cs.CV2025

NTIRE 2025 Challenge on Cross-Domain Few-Shot Object Detection: Methods and Results

Yuqian Fu, Xingyu Qiu, Bin Ren +59

Cross-Domain Few-Shot Object Detection (CD-FSOD) poses significant challenges to existing object detection and few-shot detection models when applied across domains. In conjunction…

cs.CV2025

Fractal-IR: A Unified Framework for Efficient and Scalable Image Restoration

Yawei Li, Bin Ren, Jingyun Liang +5

While vision transformers achieve significant breakthroughs in various image restoration (IR) tasks, it is still challenging to efficiently scale them across multiple types of degr…