6 papers
VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis
Shubhashis Roy Dipta, Tz-Ying Wu, Subarna Tripathi
We propose VC-Inspector, a lightweight, open-source large multimodal model (LMM) for reference-free evaluation of video captions, with a focus on factual accuracy. Unlike existing…
Search2Motion: Training-Free Object-Level Motion Control via Attention-Consensus Search
Sainan Liu, Tz-Ying Wu, Hector A Valdez +1
We present Search2Motion, a training-free framework for object-level motion editing in image-to-video generation. Unlike prior methods requiring trajectories, bounding boxes, masks…
Harnessing Object Grounding for Time-Sensitive Video Understanding
Tz-Ying Wu, Sharath Nittur Sridhar, Subarna Tripathi
We propose to improve the time-sensitive video understanding (TSV) capability of video large language models (Video-LLMs) with grounded objects (GO). We hypothesize that TSV tasks…
EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs
Ivan Rodin, Tz-Ying Wu, Kyle Min +4
We introduce EASG-Bench, a question-answering benchmark for egocentric videos where the question-answering pairs are created from spatio-temporally grounded dynamic scene graphs ca…
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models
Tz-Ying Wu, Tahani Trigui, Sharath Nittur Sridhar +2
In this paper, we introduce VideoNarrator, a novel training-free pipeline designed to generate dense video captions that offer a structured snapshot of video content. These caption…
Ego-VPA: Egocentric Video Understanding with Parameter-efficient Adaptation
Tz-Ying Wu, Kyle Min, Subarna Tripathi +1
Video understanding typically requires fine-tuning the large backbone when adapting to new domains. In this paper, we leverage the egocentric video foundation models (Ego-VFMs) bas…