8 papers
FAGER: Factually Grounded Evaluation and Refinement of Text-to-Image Models
Youngsun Lim, Cusuh Ham, Pin-Yu Chen +1
Existing text-to-image (T2I) evaluation metrics mainly assess whether generated images align with information explicitly stated in the prompt, but often fail to capture factual req…
ID-Sim: An Identity-Focused Similarity Metric
Julia Chae, Nicholas Kolkin, Jui-Hsien Wang +3
Humans have remarkable selective sensitivity to identities -- easily distinguishing between highly similar identities, even across significantly different contexts such as diverse…
TokenDial: Continuous Attribute Control in Text-to-Video via Spatiotemporal Token Offsets
Zhixuan Liu, Peter Schaldenbrand, Yijun Li +5
We present TokenDial, a framework for continuous, slider-style attribute control in pretrained text-to-video generation models. While modern generators produce strong holistic vide…
TRACE: Object Motion Editing in Videos with First-Frame Trajectory Guidance
Quynh Phung, Long Mai, Cusuh Ham +3
We study object motion path editing in videos, where the goal is to alter a target object's trajectory while preserving the original scene content. Unlike prior video editing metho…
DreamLoop: Controllable Cinemagraph Generation from a Single Photograph
Aniruddha Mahapatra, Long Mai, Cusuh Ham +1
Cinemagraphs, which combine static photographs with selective, looping motion, offer unique artistic appeal. Generating them from a single photograph in a controllable manner is pa…
CineVerse: Consistent Keyframe Synthesis for Cinematic Scene Composition
Quynh Phung, Long Mai, Fabian David Caba Heilbron +3
We present CineVerse, a novel framework for the task of cinematic scene composition. Similar to traditional multi-shot generation, our task emphasizes the need for consistency and…