3 papers
cs.CV2026
Video Generation Models are General-Purpose Vision Learners
Letian Wang, Chuhan Zhang, Rishabh Kabra +9
Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a genera…
cs.CV2026
THFM: A Unified Video Foundation Model for 4D Human Perception and Beyond
Letian Wang, Andrei Zanfir, Eduard Gabriel Bazavan +2
We present THFM, a unified video foundation model for human-centric perception that jointly addresses dense tasks (depth, normals, segmentation, dense pose) and sparse tasks (2d/3d…
cs.LG2024
Applications of fractional calculus in learned optimization
Teodor Alexandru Szente, James Harrison, Mihai Zanfir +1
Fractional gradient descent has been studied extensively, with a focus on its ability to extend traditional gradient descent methods by incorporating fractional-order derivatives.…