2 papers
cs.CV2025
LEDiT: Your Length-Extrapolatable Diffusion Transformer without Positional Encoding
Shen Zhang, Siyuan Liang, Yaning Tan +9
Diffusion transformers (DiTs) struggle to generate images at resolutions higher than their training resolutions. The primary obstacle is that the explicit positional encodings(PE),…
cs.CV2024
1st Place Solution of Multiview Egocentric Hand Tracking Challenge ECCV2024
Minqiang Zou, Zhi Lv, Riqiang Jin +4
Multi-view egocentric hand tracking is a challenging task and plays a critical role in VR interaction. In this report, we present a method that uses multi-view input images and cam…