9 papers
LoopViT: Scaling Visual ARC with Looped Transformers
Wen-Jie Shu, Xuerui Qiu, Rui-Jie Zhu +3
Recent advances in visual reasoning have leveraged vision transformers to tackle the ARC-AGI benchmark. However, we argue that the feed-forward architecture, where computational de…
Unified Personalized Understanding, Generating and Editing
Yu Zhong, Tianwei Lin, Ruike Zhu +9
Unified large multimodal models (LMMs) have achieved remarkable progress in general-purpose multimodal understanding and generation. However, they still operate under a ``one-size-…
AMS-QUANT: Adaptive Mantissa Sharing for Floating-point Quantization
Mengtao Lv, Ruiqi Zhu, Xinyu Wang +1
Large language models (LLMs) have demonstrated remarkable capabilities in various kinds of tasks, while the billion or even trillion parameters bring storage and efficiency bottlen…
Improving Joint Embedding Predictive Architecture with Diffusion Noise
Yuping Qiu, Rui Zhu, Ying-cong Chen
Self-supervised learning has become an incredibly successful method for feature learning, widely applied to many downstream tasks. It has proven especially effective for discrimina…
UMDATrack: Unified Multi-Domain Adaptive Tracking Under Adverse Weather Conditions
Siyuan Yao, Rui Zhu, Ziqi Wang +3
Visual object tracking has gained promising progress in past decades. Most of the existing approaches focus on learning target representation in well-conditioned daytime data, whil…
Unlabeled Data vs. Pre-trained Knowledge: Rethinking SSL in the Era of Large Models
Song-Lin Lv, Rui Zhu, Tong Wei +2
Semi-supervised learning (SSL) alleviates the cost of data labeling process by exploiting unlabeled data and has achieved promising results. Meanwhile, with the development of larg…