76 citations · 76 across the 4 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
TuringViT: Making SOTA Vision Transformers Accessible to All
Qiman Wu, Hanlin Chen, Lyujie Chen +19
Modern VLMs and VLA systems commonly adopt off-the-shelf ViTs such as SigLIP2 as visual encoders, but diverse downstream requirements in latency, temporal modeling, and VLM integra…
cs.CV2022★ 76 cited
RTFormer: Efficient Design for Real-Time Semantic Segmentation with Transformer
Jian Wang, Chenhui Gou, Qiman Wu +4
Recently, transformer-based networks have shown impressive results in semantic segmentation. Yet for real-time semantic segmentation, pure CNN-based approaches still dominate in th…
cs.CV2022
MixFormer: Mixing Features across Windows and Dimensions
Qiang Chen, Qiman Wu, Jian Wang +5
While local-window self-attention performs notably in vision tasks, it suffers from limited receptive field and weak modeling capability issues. This is mainly because it performs…