5 papers
On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation
Sicheng Zhang, Zhonghao Yan, Binzhu Xie +4
Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual perf…
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP
Sicheng Zhang, Muzammal Naseer, Binzhu Xie +5
CLIP and its variants are widely adopted visual backbones in multimodal systems, but their pretraining remains dominated by descriptive image-text alignment. As downstream applicat…
Uncovering What, Why and How: A Comprehensive Benchmark for Causation Understanding of Video Anomaly
Hang Du, Sicheng Zhang, Binzhu Xie +16
Video anomaly understanding (VAU) aims to automatically comprehend unusual occurrences in videos, thereby enabling various applications such as traffic surveillance and industrial…
EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning
Binzhu Xie, Shi Qiu, Sicheng Zhang +5
Robust 3D hand reconstruction in egocentric vision is challenging due to depth ambiguity, self-occlusion, and complex hand-object interactions. Prior methods mitigate these issues…
Trade-offs in Image Generation: How Do Different Dimensions Interact?
Sicheng Zhang, Binzhu Xie, Zhonghao Yan +7
Model performance in text-to-image (T2I) and image-to-image (I2I) generation often depends on multiple aspects, including quality, alignment, diversity, and robustness. However, mo…