activity
20232026
most citedCoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation

3 citations · 17 across the 26 of their papers we have counts for

collaborators
Showing 2026 · cs.CVShow all

5 papers · 2 filters

cs.CV2026

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA

Ce Zhang, Ziyang Wang, Yulu Pan +6

Grounded long-video question answering (Grounded LVQA) requires answering a question about a long video while localizing the short evidence interval that supports the answer. Recen…

cs.CV2026

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs

Yuxuan Fan, Gyusik Seo, Jing Hao +3

Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic meaning arises from deliberate combinatio…

cs.CV2026

V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising

Han Lin, Xichen Pan, Zun Wang +4

Pixel-space diffusion has recently re-emerged as a strong alternative to latent diffusion, enabling high-quality generation without pretrained autoencoders. However, standard pixel…

cs.CV2026

AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories

Zun Wang, Han Lin, Jaehong Yoon +3

Maintaining spatial world consistency over long horizons remains a central challenge for camera-controllable video generation. Existing memory-based approaches often condition gene…

cs.CV2026

Hierarchy-Aware Multimodal Unlearning for Medical AI

Fengli Wu, Vaidehi Patil, Jaehong Yoon +2

Pretrained Multimodal Large Language Models (MLLMs) are increasingly used in sensitive domains such as medical AI, where privacy regulations like HIPAA and GDPR require specific re…