5 papers
DefaultShift: Auditing Semantic Default Shift in Accelerated Text-to-Image Models
Xuanhua Yin, Chuanzhi Xu, Shunqi Mao +2
Few-step text-to-image models increasingly replace slower generators, yet acceleration can silently change distributions over unspecified attributes even when individual outputs re…
Calibrate What You SHIP: Post-Selection Risk Control for Verifier-Guided Text-to-Image Generation
Xuanhua Yin, Shunqi Mao, Wei Guo +2
Verifier-guided text-to-image systems increasingly use test-time search to select, refine, or stop among multiple candidates, yet release thresholds are often calibrated on individ…
MotionHalluc: Diagnosing Kinematic Hallucinations in Fine-Grained Motion Reasoning
Weile Guo, Shenghong He, Danying Mo +3
Motion instruction generation in cross-video comparison aims to produce corrective feedback that describes the differences between a query and a reference motion. However, existing…
Collapse of Patches: Ranking Image Patches for Efficient Visual Modeling
Wei Guo, Shunqi Mao, Zhuonan Liang +3
Observing certain patches in an image reduces the uncertainty of others. Their realization lowers the distribution entropy of each remaining patch feature, analogous to collapsing…
Gotta Hear Them All: Towards Sound Source Aware Audio Generation
Wei Guo, Heng Wang, Jianbo Ma +1
Audio synthesis has broad applications in multimedia. Recent advancements have made it possible to generate relevant audios from inputs describing an audio scene, such as images or…