4 papers
Text-Guided Visual Representation Learning for Robust Multimodal E-Commerce Recommendation
Yufei Guo, Jing Ma, Tianlu Zhang +5
Multimodal item embeddings are crucial for e-commerce item-to-item (I2I) retrieval, yet real-world product images often contain promotional overlays and background clutter that inj…
Tracking and Segmenting Anything in Any Modality
Tianlu Zhang, Qiang Zhang, Guiguang Ding +1
Tracking and segmentation play essential roles in video understanding, providing basic positional information and temporal association of objects within video sequences. Despite th…
THU-Warwick Submission for EPIC-KITCHEN Challenge 2025: Semi-Supervised Video Object Segmentation
Mingqi Gao, Haoran Duan, Tianlu Zhang +1
In this report, we describe our approach to egocentric video object segmentation. Our method combines large-scale visual pretraining from SAM2 with depth-based geometric cues to ha…
Unsupervised Patch-GAN with Targeted Patch Ranking for Fine-Grained Novelty Detection in Medical Imaging
Jingkun Chen, Guang Yang, Xiao Zhang +5
Detecting novel anomalies in medical imaging is challenging due to the limited availability of labeled data for rare abnormalities, which often display high variability and subtlet…