4 papers · 1 filter
ID-Sim: An Identity-Focused Similarity Metric
Julia Chae, Nicholas Kolkin, Jui-Hsien Wang +3
Humans have remarkable selective sensitivity to identities -- easily distinguishing between highly similar identities, even across significantly different contexts such as diverse…
TokenDial: Continuous Attribute Control for Text-to-Video Generation in Visual Dial Space
Zhixuan Liu, Peter Schaldenbrand, Yijun Li +5
In video diffusion transformers, visual patch tokens maintain explicit correspondence to space and time. We hypothesize that their channel dimension can serve as a semantic control…
Generative Timelines for Instructed Visual Assembly
Alejandro Pardo, Jui-Hsien Wang, Bernard Ghanem +3
The objective of this work is to manipulate visual timelines (e.g. a video) through natural language instructions, making complex timeline editing tasks accessible to non-expert or…
Koala: Key frame-conditioned long video-LLM
Reuben Tan, Ximeng Sun, Ping Hu +5
Long video question answering is a challenging task that involves recognizing short-term activities and reasoning about their fine-grained relationships. State-of-the-art video Lar…