63 citations · 63 across the 2 of their papers we have counts for
3 papers
cs.CL2024
Two are better than one: Context window extension with multi-grained self-injection
Wei Han, Pan Zhou, Soujanya Poria +1
The limited context window of contemporary large language models (LLMs) remains a huge barrier to their broader application across various domains. While continual pre-training on…
cs.CV2024★ 63 cited
Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
Hao Fei, Shengqiong Wu, Meishan Zhang +3
While pre-training large-scale video-language models (VLMs) has shown remarkable potential for various downstream video-language tasks, existing VLMs can still suffer from certain…
cs.CV2024
SMPLer: Taming Transformers for Monocular 3D Human Shape and Pose Estimation
Xiangyu Xu, Lijuan Liu, Shuicheng Yan
Existing Transformers for monocular 3D human shape and pose estimation typically have a quadratic computation and memory complexity with respect to the feature length, which hinder…