24 papers
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA
Ce Zhang, Ziyang Wang, Yulu Pan +6
Grounded long-video question answering (Grounded LVQA) requires answering a question about a long video while localizing the short evidence interval that supports the answer. Recen…
MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs
Yuxuan Fan, Gyusik Seo, Jing Hao +3
Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinatio…
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
Zun Wang, Jaemin Cho, Jialu Li +4
Recent approaches for video generation with camera control often create anchor videos (i.e., rendered videos that approximate desired camera motions) to guide diffusion models as a…
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
Daeun Lee, Jaehong Yoon, Jaemin Cho +1
Recent text-to-video (T2V) diffusion models have made remarkable progress in generating high-quality videos. However, they often struggle to align with complex text prompts, partic…
SiLVR: A Simple Language-based Video Reasoning Framework
Ce Zhang, Yan-Bo Lin, Ziyang Wang +2
Recent advances in test-time optimization have led to remarkable reasoning capabilities in Large Language Models (LLMs), enabling them to solve highly complex problems in math and…
V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising
Han Lin, Xichen Pan, Zun Wang +4
Pixel-space diffusion has recently re-emerged as a strong alternative to latent diffusion, enabling high-quality generation without pretrained autoencoders. However, standard pixel…