6 citations · 15 across the 13 of their papers we have counts for
9 papers · 1 filter
TokenDial: Continuous Attribute Control for Text-to-Video Generation in Visual Dial Space
Zhixuan Liu, Peter Schaldenbrand, Yijun Li +5
In video diffusion transformers, visual patch tokens maintain explicit correspondence to space and time. We hypothesize that their channel dimension can serve as a semantic control…
ShapeShift: Text-to-Mosaic Synthesis via Semantic Phase-Field Guidance
Vihaan Misra, Peter Schaldenbrand, Jean Oh
We present ShapeShift, a method for arranging rigid objects into configurations that visually convey semantic concepts specified by natural language. While pretrained diffusion mod…
SCoFT: Self-Contrastive Fine-Tuning for Equitable Image Generation
Zhixuan Liu, Peter Schaldenbrand, Beverley-Claire Okogwu +5
Accurate representation in media is known to improve the well-being of the people who consume it. Generative image models trained on large web-crawled datasets such as LAION are kn…
Robot Synesthesia: A Sound and Emotion Guided AI Painter
Vihaan Misra, Peter Schaldenbrand, Jean Oh
If a picture paints a thousand words, sound may voice a million. While recent robotic painting and image synthesis methods have achieved progress in generating visuals from text in…
Towards Equitable Representation in Text-to-Image Synthesis Models with the Cross-Cultural Understanding Benchmark (CCUB) Dataset
Zhixuan Liu, Youeun Shin, Beverley-Claire Okogwu +5
It has been shown that accurate representation in media improves the well-being of the people who consume it. By contrast, inaccurate representations can negatively affect viewers…
Towards Real-Time Text2Video via CLIP-Guided, Pixel-Level Optimization
Peter Schaldenbrand, Zhixuan Liu, Jean Oh
We introduce an approach to generating videos based on a series of given language descriptions. Frames of the video are generated sequentially and optimized by guidance from the CL…