2 citations · 2 across the 4 of their papers we have counts for
3 papers · 1 filter
SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos
Jinlin Wu, Felix Holm, Chuxi Chen +17
While foundation models have advanced surgical video analysis, current approaches rely predominantly on pixel-level reconstruction objectives that waste model capacity on low-level…
Compositional Inversion for Stable Diffusion Models
Xulu Zhang, Xiao-Yong Wei, Jinlin Wu +4
Inversion methods, such as Textual Inversion, generate personalized images by incorporating concepts of interest provided by user images. However, existing methods often suffer fro…
FRCSyn Challenge at WACV 2024:Face Recognition Challenge in the Era of Synthetic Data
Pietro Melzi, Ruben Tolosana, Ruben Vera-Rodriguez +44
Despite the widespread adoption of face recognition technology around the world, and its remarkable performance on current benchmarks, there are still several challenges that must…