346 citations
- Shenzhen UniversityCN20 papers
- Ping An (China)9 papers
- Chinese Academy of SciencesCN8 papers
- Tsinghua UniversityCN7 papers
- University of Science and Technology of ChinaCN7 papers
- Peking UniversityCN6 papers
- University of Chinese Academy of SciencesCN6 papers
- Xi'an Jiaotong UniversityCN6 papers
- Beijing University of Posts and TelecommunicationsCN5 papers
- City University of Hong KongHK5 papers
- Guangzhou UniversityCN5 papers
- Harbin Institute of TechnologyCN5 papers
11 papers · 1 filter
A3-FPN: Asymptotic Content-Aware Pyramid Attention Network for Dense Visual Prediction
Meng'en Qin, Yu Song, Quanling Zhao +3
Learning multi-scale representations is the common strategy to tackle object scale variation in dense prediction tasks. Although existing feature pyramid networks have greatly adva…
WeatherRemover: All-in-one Adverse Weather Removal with Multi-scale Feature Map Compression
Weikai Qu, Sijun Liang, Cheng Pan +6
Photographs taken in adverse weather conditions often suffer from blurriness, occlusion, and low brightness due to interference from rain, snow, and fog. These issues can significa…
VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
Qiang Wang, Xinyuan Gao, Yuhang He +5
Existing Video Detailed Captioning (VDC) methods predominantly rely on costly human annotations or distillation from powerful proprietary models, creating a dependency on external…
Endoscopic Depth Estimation Based on Deep Learning: A Survey
Ke Niu, Zeyun Liu, Xue Feng +5
Endoscopic depth estimation is a critical technology for improving the safety and precision of minimally invasive surgery. It has attracted considerable attention from researchers…
Taming Anomalies with Down-Up Sampling Networks: Group Center Preserving Reconstruction for 3D Anomaly Detection
Hanzhe Liang, Jie Zhang, Tao Dai +3
Reconstruction-based methods have demonstrated very promising results for 3D anomaly detection. However, these methods face great challenges in handling high-precision point clouds…
Unleashing the Potential of All Test Samples: Mean-Shift Guided Test-Time Adaptation
Jizhou Han, Chenhao Ding, SongLin Dong +3
Visual-language models (VLMs) like CLIP exhibit strong generalization but struggle with distribution shifts at test time. Existing training-free test-time adaptation (TTA) methods…