7 papers · 1 filter
Dr. Seg: Revisiting GRPO Training for Visual Large Language Models through Perception-Oriented Design
Haoxiang Sun, Tao Wang, Chenwei Tang +2
Following the success of Group Relative Policy Optimization (GRPO) in foundation LLMs, an increasing number of works have sought to adapt GRPO to Visual Large Language Models (VLLM…
Adaptive Image Zoom-in with Bounding Box Transformation for UAV Object Detection
Tao Wang, Chenyu Lin, Chenwei Tang +5
Detecting objects from UAV-captured images is challenging due to the small object size. In this work, a simple and efficient adaptive zoom-in framework is explored for object detec…
Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval
Yifan Wang, Tao Wang, Chenwei Tang +5
Recently, prompt learning has demonstrated remarkable success in adapting pre-trained Vision-Language Models (VLMs) to various downstream tasks such as image classification. Howeve…
EBAD-Gaussian: Event-driven Bundle Adjusted Deblur Gaussian Splatting
Yufei Deng, Yuanjian Wang, Rong Xiao +6
While 3D Gaussian Splatting (3D-GS) achieves photorealistic novel view synthesis, its performance degrades with motion blur. In scenarios with rapid motion or low-light conditions,…
SaENeRF: Suppressing Artifacts in Event-based Neural Radiance Fields
Yuanjian Wang, Yufei Deng, Rong Xiao +4
Event cameras are neuromorphic vision sensors that asynchronously capture changes in logarithmic brightness changes, offering significant advantages such as low latency, low power…
Memory-Augmented Dual-Decoder Networks for Multi-Class Unsupervised Anomaly Detection
Jingyu Xing, Chenwei Tang, Tao Wang +5
Recent advances in unsupervised anomaly detection (UAD) have shifted from single-class to multi-class scenarios. In such complex contexts, the increasing pattern diversity has brou…