collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

SPAE: Spectrally Guided Autoencoder for Pretrained Visual Latents

Yibin Huang, Jixiang Hong, Zongzhao Li +8

Latents from vision foundation models (VFMs) are semantically rich and well suited for visual understanding. Recent representation autoencoder methods such as RAE have shown that t…

cs.CV2026

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization

Tianhong Zhou, Mingyang Han, Boyu Li +8

Audio-visual feature extraction is a fundamental component of multimodal understanding and generation tasks. However, existing evaluation protocols for feature extraction models ex…

cs.CV2026

Deep Pre-Alignment for VLMs

Tianyu Yu, Kechen Fang, Zihao Wan +5

Most Vision Language Models (VLMs) directly map outputs from ViT encoders to the LLM via a lightweight projector. While effective, recent analysis suggests this architecture suffer…

cs.CV2026

DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing

Hanqing Yang, Qiang Zhou, Yongchao Du +6

Recent image editing models have achieved strong visual fidelity but often struggle with tasks requiring complex reasoning. To investigate and enhance the reasoning-grounded planni…

cs.CV2026

ViT: Unlocking Test-Time Training in Vision

Dongchen Han, Yining Li, Tianyu Li +6

Test-Time Training (TTT) has recently emerged as a promising direction for efficient sequence modeling. TTT reformulates attention operation as an online learning problem, construc…

cs.CV2026

AdaGen: Learning Adaptive Policy for Image Synthesis

Zanlin Ni, Yulin Wang, Yeguo Hua +5

Recent advances in image synthesis have been propelled by powerful generative models, such as Masked Generative Transformers (MaskGIT), autoregressive models, diffusion models, and…