activity
20242026
collaborators

5 papers

cs.CV2026

VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning

Zhong-Yu Li, Ruoyi Du, Juncheng Yan +6

Recent progress in diffusion models significantly advances various image generation tasks. However, the current mainstream approach remains focused on building task-specific models…

cs.CV2025

Enhancing Representations through Heterogeneous Self-Supervised Learning

Zhong-Yu Li, Bo-Wen Yin, Yongxiang Liu +2

Incorporating heterogeneous representations from different architectures has facilitated various vision tasks, e.g., some hybrid networks combine transformers and convolutions. How…

cs.CV2024

Multi-Token Enhancing for Vision Representation Learning

Zhong-Yu Li, Yu-Song Hu, Bo-Wen Yin +1

Vision representation learning, especially self-supervised learning, is pivotal for various vision applications. Ensemble learning has also succeeded in enhancing the performance a…

cs.CV2024

PR-MIM: Delving Deeper into Partial Reconstruction in Masked Image Modeling

Zhong-Yu Li, Yunheng Li, Deng-Ping Fan +1

Masked image modeling has achieved great success in learning representations but is limited by the huge computational costs. One cost-saving strategy makes the decoder reconstruct…

cs.CV2024

Towards RAW Object Detection in Diverse Conditions

Zhong-Yu Li, Xin Jin, Boyuan Sun +2

Existing object detection methods often consider sRGB input, which was compressed from RAW data using ISP originally designed for visualization. However, such compression might los…