2 papers
cs.CV2026
What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
Chen-Yi Lu, Yueh-Shao Chen, Somali Chaterji
Contrastive vision-language models such as CLIP map semantically opposite phrases (e.g., "a dog" vs. "not a dog") to nearly identical embeddings, rendering them insensitive to nega…
cs.CV2025
SKALD: Learning-Based Shot Assembly for Coherent Multi-Shot Video Creation
Chen Yi Lu, Md Mehrab Tanjim, Ishita Dasgupta +4
We present SKALD, a multi-shot video assembly method that constructs coherent video sequences from candidate shots with minimal reliance on text. Central to our approach is the Lea…