4 papers
Self-supervised pretraining for an iterative image size agnostic vision transformer
Nedyalko Prisadnikov, Danda Pani Paudel, Yuqian Fu +1
Vision Transformers (ViTs) dominate self-supervised learning (SSL). While they have proven highly effective for large-scale pretraining, they are computationally inefficient and sc…
StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
Yifei Wang, Zhenkai Li, Tianwen Qian +4
As embodied intelligence advances toward real-world deployment, the ability to continuously perceive and reason over streaming visual inputs becomes essential. In such settings, an…
ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
Qi'ao Xu, Tianwen Qian, Yuqian Fu +5
A core capability towards general embodied intelligence lies in localizing task-relevant objects from an egocentric perspective, formulated as Spatio-Temporal Video Grounding (STVG…
MAT: Multi-Range Attention Transformer for Efficient Image Super-Resolution
Chengxing Xie, Xiaoming Zhang, Linze Li +4
Image super-resolution (SR) has significantly advanced through the adoption of Transformer architectures. However, conventional techniques aimed at enlarging the self-attention win…