activity
20242026
collaborators

5 papers

cs.CV2026

DefaultShift: Auditing Semantic Default Shift in Accelerated Text-to-Image Models

Xuanhua Yin, Chuanzhi Xu, Shunqi Mao +2

Few-step text-to-image models increasingly replace slower generators, yet acceleration can silently change distributions over unspecified attributes even when individual outputs re…

cs.CV2026

Calibrate What You SHIP: Post-Selection Risk Control for Verifier-Guided Text-to-Image Generation

Xuanhua Yin, Shunqi Mao, Wei Guo +2

Verifier-guided text-to-image systems increasingly use test-time search to select, refine, or stop among multiple candidates, yet release thresholds are often calibrated on individ…

cs.CV2026

MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning

Weile Guo, Shenghong He, Danying Mo +3

Motion instruction generation in cross-video comparison aims to produce corrective feedback that describes the differences between a query and a reference motion. However, existing…

cs.CV2025

Collapse of Patches: Ranking Image Patches for Efficient Visual Modeling

Wei Guo, Shunqi Mao, Zhuonan Liang +3

Observing certain patches in an image reduces the uncertainty of others. Their realization lowers the distribution entropy of each remaining patch feature, analogous to collapsing…

cs.MM2024

Gotta Hear Them All: Towards Sound Source Aware Audio Generation

Wei Guo, Heng Wang, Jianbo Ma +1

Audio synthesis has broad applications in multimedia. Recent advancements have made it possible to generate relevant audios from inputs describing an audio scene, such as images or…