4 papers
Search2Motion: Training-Free Object-Level Motion Control via Attention-Consensus Search
Sainan Liu, Tz-Ying Wu, Hector A Valdez +1
We present Search2Motion, a training-free framework for object-level motion editing in image-to-video generation. Unlike prior methods requiring trajectories, bounding boxes, masks…
SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video
Hector A. Valdez, Kyle Min, Subarna Tripathi
Pretraining egocentric vision-language models has become essential to improving downstream egocentric video-text tasks. These egocentric foundation models commonly use the transfor…
Contrastive Language Video Time Pre-training
Hengyue Liu, Kyle Min, Hector A. Valdez +1
We introduce LAVITI, a novel approach to learning language, video, and temporal representations in long-form videos via contrastive learning. Different from pre-training on video-t…
Positional-Unigram Byte Models for Generalized TLS Fingerprinting
Hector A. Valdez, Sean McPherson
We use positional-unigram byte models along with maximum likelihood for generalized TLS fingerprinting and empirically show that it is robust to cipher stunting. Our approach creat…