activity
20182024
most citedStyleT2I: Toward Compositional and High-Fidelity Text-to-Image Synthesis

2 citations · 3 across the 4 of their papers we have counts for

collaborators

8 papers

cs.CV2024

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation

Hao Li, Shamit Lal, Zhiheng Li +9

We empirically study the scaling properties of various Diffusion Transformers (DiTs) for text-to-image generation by performing extensive and rigorous ablations, including training…

cs.CV20241 cited

FairRAG: Fair Human Generation via Fair Retrieval Augmentation

Robik Shrestha, Yang Zou, Qiuyu Chen +3

Existing text-to-image generative models reflect or even amplify societal biases ingrained in their training data. This is especially concerning for human image generation where mo…

cs.CV2024

Discover and Mitigate Multiple Biased Subgroups in Image Classifiers

Zeliang Zhang, Mingqian Feng, Zhiheng Li +1

Machine learning models can perform well on in-distribution data but often fail on biased subgroups that are underrepresented in the training data, hindering the robustness of mode…

cs.CV20222 cited

StyleT2I: Toward Compositional and High-Fidelity Text-to-Image Synthesis

Zhiheng Li, Martin Renqiang Min, Kai Li +1

Although progress has been made for text-to-image synthesis, previous methods fall short of generalizing to unseen or underrepresented attribute compositions in the input text. Lac…

cs.CV2021

Discover the Unknown Biased Attribute of an Image Classifier

Zhiheng Li, Chenliang Xu

Recent works find that AI algorithms learn biases from data. Therefore, it is urgent and vital to identify biases in AI algorithms. However, the previous bias identification pipeli…

cs.CV2020

Learning a Weakly-Supervised Video Actor-Action Segmentation Model with a Wise Selection

Jie Chen, Zhiheng Li, Jiebo Luo +1

We address weakly-supervised video actor-action segmentation (VAAS), which extends general video object segmentation (VOS) to additionally consider action labels of the actors. The…