collaborators

10 papers

cs.CV2026

SPAE: Spectrally Guided Autoencoder for Pretrained Visual Latents

Yibin Huang, Jixiang Hong, Zongzhao Li +8

Latents from vision foundation models (VFMs) are semantically rich and well suited for visual understanding. Recent representation autoencoder methods such as RAE have shown that t…

cs.CV2026

AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization

Tianhong Zhou, Mingyang Han, Boyu Li +8

Audio-visual feature extraction is a fundamental component of multimodal understanding and generation tasks. However, existing evaluation protocols for feature extraction models ex…

cs.SD2026

FreeSonic: Training-Free Temporal-Aware Decoupled Attention for Precise Audio Editing

Yuxuan Jiang, Mingyang Han, Yusheng Dai +12

Text-to-audio (TTA) generation has made significant strides, yet achieving precise and consistent audio editing remains a major challenge. However, existing methods struggle to bal…

cs.CL2026

The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language Models

Zanlin Ni, Shenzhi Wang, Yang Yue +8

Diffusion Large Language Models (dLLMs) break the rigid left-to-right constraint of traditional LLMs, enabling token generation in arbitrary orders. Intuitively, this flexibility i…

cs.CV2026

Deep Pre-Alignment for VLMs

Tianyu Yu, Kechen Fang, Zihao Wan +5

Most Vision Language Models (VLMs) directly map outputs from ViT encoders to the LLM via a lightweight projector. While effective, recent analysis suggests this architecture suffer…

cs.CV2026

DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing

Hanqing Yang, Qiang Zhou, Yongchao Du +6

Recent image editing models have achieved strong visual fidelity but often struggle with tasks requiring complex reasoning. To investigate and enhance the reasoning-grounded planni…