4 papers
Posterior Augmented Flow Matching
George Stoica, Sayak Paul, Matthew Wallingford +6
Flow matching (FM) trains a time-dependent vector field that transports samples from a simple prior to a complex data distribution. However, for high-dimensional images, each train…
Resolving Interference (RI): Disentangling Models for Improved Model Merging
Pratik Ramesh, George Stoica, Arun Iyer +2
Model merging has shown that multitask models can be created by directly combining the parameters of different models that are each specialized on tasks of interest. However, model…
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
Christopher Clark, Jieyu Zhang, Zixian Ma +18
Today's strongest video-language models (VLMs) remain proprietary. The strongest open-weight models either rely on synthetic data from proprietary VLMs, effectively distilling from…
Structure From Tracking: Distilling Structure-Preserving Motion for Video Generation
Yang Fei, George Stoica, Jingyuan Liu +4
Reality is a dance between rigid constraints and deformable structures. For video models, that means generating motion that preserves fidelity as well as structure. Despite progres…