Publications (38)
Vision Transformers Need More Than Registers
Cheng Shi, Yizhou Yu, Sibei Yang
Vision Transformers (ViTs), when pre-trained on large-scale data, provide general-purpose representations for diverse downstream tasks. However, artifacts in ViTs are widely observ…
Joint Graph Rewiring and Feature Denoising via Spectral Resonance
Jonas Linkerhägner, Cheng Shi, Ivan DokmaniÄ
When learning from graph data, the graph and the node features both give noisy information about the node labels. In this paper we propose an algorithm to jointly denoise the featu…
Free-Bloom: Zero-Shot Text-to-Video Generator with LLM Director and LDM Animator
Hanzhuo Huang, Yufan Feng, Cheng Shi +3
Text-to-video is a rapidly growing research area that aims to generate a semantic, identical, and temporal coherence sequence of frames that accurately align with the input text pr…
Weighted Community Detection and Data Clustering Using Message Passing
Cheng Shi, Yanchen Liu, Pan Zhang
Grouping objects into clusters based on similarities or weights between them is one of the most important problems in science and engineering. In this work, by extending message pa…
EdaDet: Open-Vocabulary Object Detection Using Early Dense Alignment
Cheng Shi, Sibei Yang
Vision-language models such as CLIP have boosted the performance of open-vocabulary object detection, where the detector is trained on base categories but required to detect novel…
Pretraining a Foundation Model for Small-Molecule Natural Products
Yuheng Ding, Bo Qiang, Shaoning Li +8
Natural products, as metabolites from microorganisms, animals, or plants, exhibit diverse biological activities, making them crucial for drug discovery. Nowadays, existing deep lea…