2 papers
cs.CV2026
To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs
Rui Hong, Shuxue Quan
When VLMs answer correctly, do they genuinely rely on visual information? We introduce a Tri-Layer Diagnostic Framework with three per-sample metrics: Latent Anomaly Detection, Vis…
cs.CV2026
Motion-Adaptive Temporal Attention for Lightweight Video Generation with Stable Diffusion
Rui Hong, Shuxue Quan
We present a motion-adaptive temporal attention mechanism for parameter-efficient video generation built upon frozen Stable Diffusion models. Rather than treating all video content…