1 citations · 1 across the 4 of their papers we have counts for
8 papers · 1 filter
GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
Amita Kamath, Kai-Wei Chang, Ranjay Krishna +3
Automating Text-to-Image (T2I) model evaluation is challenging; a judge model must be used to score correctness, and test prompts must be selected to be challenging for current T2I…
Structure From Tracking: Distilling Structure-Preserving Motion for Video Generation
Yang Fei, George Stoica, Jingyuan Liu +4
Reality is a dance between rigid constraints and deformable structures. For video models, that means generating motion that preserves fidelity as well as structure. Despite progres…
Visual Representations inside the Language Model
Benlin Liu, Amita Kamath, Madeleine Grunde-McLaughlin +2
Despite interpretability work analyzing VIT encoders and transformer activations, we don't yet understand why Multimodal Language Models (MLMs) struggle on perception-heavy tasks.…
Reinforced Visual Perception with Tools
Zetong Zhou, Dongping Chen, Zixian Ma +6
Visual reasoning, a cornerstone of human intelligence, encompasses complex perceptual and logical processes essential for solving diverse visual problems. While advances in compute…
MultiRef: Controllable Image Generation with Multiple Visual References
Ruoxi Chen, Dongping Chen, Siyuan Wu +6
Visual designers naturally draw inspiration from multiple visual references, combining diverse elements and aesthetic principles to create artwork. However, current image generativ…
RefTok: Reference-Based Tokenization for Video Generation
Xiang Fan, Xiaohang Sun, Kushan Thakkar +4
Effectively handling temporal redundancy remains a key challenge in learning video models. Prevailing approaches often treat each set of frames independently, failing to effectivel…