collaborators

5 papers

cs.CL2026

On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation

Sicheng Zhang, Zhonghao Yan, Binzhu Xie +4

Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual perf…

cs.CV2026

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP

Sicheng Zhang, Muzammal Naseer, Binzhu Xie +5

CLIP and its variants are widely adopted visual backbones in multimodal systems, but their pretraining remains dominated by descriptive image-text alignment. As downstream applicat…

cs.CV2026

Uncovering What, Why and How: A Comprehensive Benchmark for Causation Understanding of Video Anomaly

Hang Du, Sicheng Zhang, Binzhu Xie +16

Video anomaly understanding (VAU) aims to automatically comprehend unusual occurrences in videos, thereby enabling various applications such as traffic surveillance and industrial…

cs.CV2026

EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning

Binzhu Xie, Shi Qiu, Sicheng Zhang +5

Robust 3D hand reconstruction in egocentric vision is challenging due to depth ambiguity, self-occlusion, and complex hand-object interactions. Prior methods mitigate these issues…

cs.CV2025

Trade-offs in Image Generation: How Do Different Dimensions Interact?

Sicheng Zhang, Binzhu Xie, Zhonghao Yan +7

Model performance in text-to-image (T2I) and image-to-image (I2I) generation often depends on multiple aspects, including quality, alignment, diversity, and robustness. However, mo…