From the 5 of 91 papers with an AI index.
17 citations
- Tsinghua UniversityCN40 papers
- Peking UniversityCN37 papers
- Institute of Modern PhysicsCN36 papers
- Carnegie Mellon UniversityUS34 papers
- Istituto Nazionale di Fisica Nucleare, Laboratori Nazionali di FrascatiIT34 papers
- Istituto Nazionale di Fisica Nucleare, Sezione di PerugiaIT34 papers
- Nanjing Normal UniversityCN34 papers
- National Centre for Nuclear ResearchPL34 papers
- South China Normal UniversityCN34 papers
- Università degli Studi del Piemonte Orientale “Amedeo Avogadro”IT34 papers
- University of BristolGB34 papers
- University of PerugiaIT34 papers
15 papers · 1 filter
SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference
Shanghao Liu, Renze Chen, Size Zheng +4
Video diffusion transformers (vDiTs) generate high quality but pay quadratic self-attention cost, making inference prohibitive at video-token scales. The challenge is input-adaptiv…
LDFE: Laplacian Decoupled Feature Enhancement Block for Dual-Stream CNN-based RGB-IR Object Detection
Wenhao Dong, Xiaoyan Luo, Linlin Yang +4
The complementary information between RGB and IR images can significantly enhance object detection performance under extreme conditions. Existing methods prefer dual-stream CNN bac…
Benchmarking the Robustness of Autonomous Driving to Environmental Illusions: A Lane Perception Perspective
Tianyuan Zhang, Xianglong Liu, Aishan Liu +6
Environmental illusions (eg., shadows, reflections, and tire marks) are naturally existing yet overlooked phenomena in real-world driving environments. They can disturb visual perc…
State Space Models Meet Remote Sensing: A Survey
Qinzhe Yang, Chenyang Liu, Jia Xu +2
State Space Models (SSMs), designed for long-range modeling, offer linear computational complexity and strong capabilities in capturing long-range dependencies. In the field of rem…
Efficient Remote Sensing Instance Segmentation with Linear-Time State Space Distilled Visual Foundation Models
Qinzhe Yang, Keyan Chen, Jia Xu +2
The computational complexity of Transformers scales quadratically with the number of tokens, which significantly constrains the efficiency of vision models, particularly recent ViT…
Prompt-Calibrated SAM 3 for Open-Vocabulary Remote Sensing Semantic Segmentation
Yanghui Song, Nanqing Liu, Haonan Yin +3
Open-vocabulary semantic segmentation (OVSS) in remote sensing images aims to segment categories beyond a fixed label space. Recent SAM 3-based methods provide a promising training…