7 papers · 1 filter
SPAE: Spectrally Guided Autoencoder for Pretrained Visual Latents
Yibin Huang, Jixiang Hong, Zongzhao Li +8
Latents from vision foundation models (VFMs) are semantically rich and well suited for visual understanding. Recent representation autoencoder methods such as RAE have shown that t…
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization
Tianhong Zhou, Mingyang Han, Boyu Li +8
Audio-visual feature extraction is a fundamental component of multimodal understanding and generation tasks. However, existing evaluation protocols for feature extraction models ex…
Deep Pre-Alignment for VLMs
Tianyu Yu, Kechen Fang, Zihao Wan +5
Most Vision Language Models (VLMs) directly map outputs from ViT encoders to the LLM via a lightweight projector. While effective, recent analysis suggests this architecture suffer…
DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing
Hanqing Yang, Qiang Zhou, Yongchao Du +6
Recent image editing models have achieved strong visual fidelity but often struggle with tasks requiring complex reasoning. To investigate and enhance the reasoning-grounded planni…
ViT: Unlocking Test-Time Training in Vision
Dongchen Han, Yining Li, Tianyu Li +6
Test-Time Training (TTT) has recently emerged as a promising direction for efficient sequence modeling. TTT reformulates attention operation as an online learning problem, construc…
AdaGen: Learning Adaptive Policy for Image Synthesis
Zanlin Ni, Yulin Wang, Yeguo Hua +5
Recent advances in image synthesis have been propelled by powerful generative models, such as Masked Generative Transformers (MaskGIT), autoregressive models, diffusion models, and…